VLDB 2026 Research / reviewers in the wild / expert
Youngjae Kim 0001
dblp:19/5848
· DBLP profile ↗
76ranked-venue papers
10as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 67 · 8 first-author · 25 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VeX: Scaling HNSW-Based Vector Search with DPU Memory and Parallelism
Hyungsun Yoo, Woojung Kim, Donghyun Min, Myungcheol Lee, Jihoon Yang, Weikuan Yu, Youngjae Kim 0001 |
CCGrid | 8 |
| 2026 | BucketLSM: Breaking the Compaction Scalability Barrier in LSM-Based Key-Value Stores
Jaewan Park, Kyungwook Min, Sungjin Byeon, Taewan Noh, Hyungi Park, Xubin He, Hong-Yeon Kim, Youngjae Kim 0001 |
CCGrid | 8 |
| 2026 | Dual-Blade: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
Bodon Jeong, Hongsu Byun, Youngjae Kim 0001, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park |
ICDCS | 3 |
| 2026 | ColdMap: Compaction-Aware Cost-Benefit Zone Cleaning for ZNS-Based Key-Value Stores
Sungjin Byeon, Kyungwook Min, Jaewan Park, Hong-Yeon Kim, Junyoung Han, Jooyoung Hwang, Zhichao Cao 0002, Youngjae Kim 0001 |
ICS | 9 |
| 2025 | Cost-Efficient VM Selection for Cloud-Based LLM Inference with KV Cache OffloadingabstractLLM inference is essential for applications like text summarization, translation, and data analysis, but the high cost of GPU instances from Cloud Service Providers (CSPs) like AWS is a major burden. This paper proposes InferSave, a cost-efficient VM selection framework for cloud-based LLM inference. InferSave optimizes KV cache offloading based on Service Level Objectives (SLOs) and workload characteristics, estimating GPU memory needs, and recommending cost-effective VM instances. Additionally, the Compute Time Calibration Function (CTCF) improves instance selection accuracy by adjusting for discrepancies between theoretical and actual GPU performance. Experiments on AWS GPU instances show that selecting lower-cost instances without KV cache offloading improves cost efficiency by up to 73.7% for online workloads, while KV cache offloading saves up to 20.19% for offline workloads. Hyunsun Chung, Myung-Hoon Cha, Hong-Yeon Kim, Youngjae Kim 0001 |
CLOUD | 6 |
| 2025 | Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG SystemsabstractRecent large language models (LLMs) face increasing inference latency as input context length and model size grow. Retrieval-augmented generation (RAG) exacerbates this by significantly increasing input tokens, leading to higher computational overhead during the prefill stage and prolonged time-to-first-token (TTFT). To address this, the paper proposes using a disk-based key-value (KV) cache to reduce the prefill computational burden, thereby shortening TTFT. We also introduce a disk-based shared KV cache management system, called Shared RAG-DCache, for multi-instance LLM RAG service environments. This system leverages the locality of documents related to user queries in RAG and queueing delays in LLM inference services to proactively generate and store disk KV caches for query-related documents, sharing them across multiple LLM instances to enhance inference performance. In experiments on a single host with 2 GPUs and 1 CPU, Shared RAG-DCache achieved a 15–71 % increase in throughput and up to a 12–65 % reduction in latency, depending on the resource configuration. Hyungwoo Lee, Jungmin So, Myung-Hoon Cha, Hong-Yeon Kim, James J. Kim, Youngjae Kim 0001 |
CLOUD | 8 |
| 2025 | ECO-KVS: Energy-Aware Compaction Offloading Mechanism for LSM-Tree Based Key-Value Stores in Edge FederationabstractIn recent years, the rise in energy consumption across infrastructure has highlighted the need for more energy-efficient technologies. This is particularly critical in edge computing environments, where resources and power are limited. Consequently, there is increasing interest in improving the energy efficiency of resource-intensive tasks on edge servers. Edge servers commonly use Log-Structured Merge-tree-based Key-Value Store (LSM-KVS), to manage continuous data streams from edge devices. A key operation in LSM-KVS, known as compaction, merges key-value pairs in a CPU-intensive and energy-demanding process. Additionally, delays during compaction can cause write stalls, blocking I/O operations and degrading performance. This creates a significant challenge in balancing energy consumption and system performance. To address these challenges, we propose ECO-KVS, a solution that improves both energy efficiency and performance in LSM-KVS by offloading compaction tasks across edge servers in an edge federation. ECO-KVS leverages a real-time learning model to predict compaction time and energy consumption, reducing write stalls and enhancing overall energy efficiency. Implemented on RocksDB, ECO-KVS achieves up to 21% higher throughput compared to the baseline RocksDB and improves the performance-to-energy efficiency ratio by up to 18 % compared to EdgePilot, a state-of-the-art solution for edge environments. Jeeseob Kim, Hongsu Byun, Myoungjoon Kim, Youngjae Kim 0001, Zaipeng Xie, Sungyong Park |
CCGrid | 5 |
| 2025 | MEMORYBRIDGE: Leveraging Cloud Resource Characteristics for Cost-Efficient Disk-Based GNN Training via Two-Level ArchitectureabstractGraph Neural Networks (GNNs) are machine learning models that process graph-structured data by learning relationships between vertices and edges, as well as graph-level characteristics. Recently, with the emergence of large graph datasets on a TB scale, dataset sizes have exceeded the memory capacity of single machines. As a result, traditional methods that load all graph data into memory have become unusable, leading to the emergence of disk-based GNN training that uses storage as a memory extension. Recent research has focused on reducing disk I/O bottlenecks in disk-based GNNs. However, disk-based GNNs face new challenges in cloud environments due to two main characteristics. First, compared to node-local environments, the significantly slower cloud storage I/O speed becomes the main bottleneck of the entire training process. Second, pre-defined virtual machines prevent users from freely utilizing desired memory sizes, bandwidth, and the latest GPU technologies. These limitations have made existing disk-based GNN research unusable in cloud environments. To overcome this, we propose MEMORYBRIDGE, a system that cost-effectively accelerates GNN training in cloud environments through a novel two-level architecture that utilizes affordable GPU resources as training nodes and remote memory resources without GPUs as memory nodes, instead of using a single expensive GPU resource. This architecture consists of two key components: (i) a mathematical solver that recommends the most cost-effective resource combination, and (ii) a cloud-specialized GNN framework that implements graphaware fixed caching and batch pipelining optimization. The experimental results show that MEMORYBRIDGE achieved a speed improvement of up to 32.7x compared to existing GNN training frameworks and a cost efficiency of 9.9x compared to alternative resource configuration strategies, effectively handling the unique problems that arise from the combination of cloud environments and GNN training. Yoochan Kim, Weikuan Yu, Hong-Yeon Kim, Youngjae Kim 0001 |
CCGrid | 4 |
| 2025 | CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector DatabasesabstractEmbedding models capture both semantic and syntactic structures of queries, often mapping different queries to similar regions in vector space. This results in nonuniform cluster access patterns in modern disk-based vector databases. While existing approaches optimize individual queries, they overlook the impact of cluster access patterns, failing to account for the locality effects of queries that access similar clusters. This oversight increases cache miss penalty. To minimize the cache miss penalty, we propose CALL, a context-aware query grouping mechanism that organizes queries based on shared cluster access patterns. Additionally, CALL incorporates a group-aware prefetching method to minimize cache misses during transitions between query groups and latency-aware cluster loading. Experimental results show that CALL reduces the 99th percentile tail latency by up to 33 % while consistently maintaining a higher cache hit ratio, substantially reducing search latency. Yeonwoo Jeong, Hyunji Cho, Kyuli Park, Youngjae Kim 0001, Sungyong Park |
HiPC | 4 |
| 2025 | ByteExpress: A High-Performance and Traffic-Efficient Inline Transfer of Small Payloads over NVMeabstractRecent computational storage devices enable host-side tasks such as SQL filtering and key-value operations to be offloaded to the device. However, these tasks often involve small payloads, typically a few dozen to hundreds of bytes, which are inefficiently handled by the conventional NVMe protocol due to its page-based DMA mechanism. Even tiny payloads incur 4 KB PCIe transfers, leading to severe bandwidth waste and increased latency. Prior approaches either break NVMe compatibility or are only effective for very small payloads on the order of a few dozen bytes. This paper presents ByteExpress, a new mechanism that efficiently transmits small payloads by placing them inline in 64-byte chunks directly into the NVMe submission queue, immediately following the NVMe command. ByteExpress requires only slight modifications to the NVMe driver and controller logic, while preserving full compatibility with existing APIs and SSD architectures. We implemented ByteExpress on the Linux NVMe driver and OpenSSD, demonstrating up to 98% reduction in PCIe traffic and 40% and 39% lower latency compared to PRP and a state-of-the-art approach, respectively, for sub-page payloads. Junhyeok Park 0002, Junghee Lee 0004, Youngjae Kim 0001 |
HotStorage | 3 |
| 2025 | DEDUPKV: A Space-Efficient and High-Performance Key-Value Store via Fine-Grained DeduplicationabstractLog-Structured Merge Tree (LSM-tree) based key-value stores excel in write-intensive environments but suffer from data duplication, consuming up to 49% of storage space in LSMtree-based key-value store deployments.Traditional solutions like compression and coarse-grained file system-level deduplication introduce overhead or have limited effectiveness.In this study, we propose DedupKV, a fine-grained deduplication framework tailored for LSM-tree, maximizing data reduction efficiency while minimizing write stalls and read overheads.DedupKV features three key innovations:(1) FLUSH-integrated inline deduplication, which removes duplicates during memory-to-storage writes; (2) WAL file-based offline deduplication, repurposing write-ahead logs to avoid double writes; and (3) elastic execution, dynamically balancing inline and offline deduplication based on memory pressure and workload intensity.Additionally, dynamic granularity management reduces deduplication metadata overhead.We implemented these four ideas in RocksDB for the first time and conducted experiments in a Linux environment.Our evaluation shows that WAL file-based offline deduplication and DedupKV outperform BlobDB by 33% and 23%, respectively, in write-heavy workloads, while reducing write amplification by 1.2×, 2×, and 1.6× for real KV datasets. Safdar Jamil, Awais Khan 0002, Xubin He, Youngjae Kim 0001 |
ICS | 4 |
| 2025 | KVACCEL: A Novel Write Accelerator for LSM-Tree-Based KV Stores with Host-SSD CollaborationabstractLog-Structured Merge (LSM) tree-based Key-Value Stores (KVSs) are widely adopted for their high performance in write-intensive environments, but they often face performance degradation due to write stalls during compaction. Prior solutions, such as regulating I/O traffic or using multiple compaction threads, can cause unexpected drops in throughput or increase host CPU usage, while hardware-based approaches using FPGA, GPU, and DPU aimed at reducing compaction duration introduce additional hardware costs. In this study, we propose KVACCEL, a novel hardware-software co-design framework that bypasses write stalls by leveraging a dual-interface SSD. KVACCEL allocates logical NAND flash space to support both block and key-value interfaces, using the key-value interface as a temporary write buffer during write stalls. This strategy significantly reduces write stalls, optimizes resource usage, and ensures consistency between the host and device by implementing an in-device LSM-based write buffer with an iterator-based range scan mechanism. Our extensive evaluation shows that for write-intensive workloads, KVACCEL outperforms ADOC by up to 17 % in terms of throughput and performance-to-CPU-utilization efficiency. For mixed read-write workloads, both demonstrate comparable performance. Hyunsun Chung, Seonghoon Ahn, Junhyeok Park 0002, Safdar Jamil, Hongsu Byun, Myungcheol Lee, Jinchun Choi, Youngjae Kim 0001 |
IPDPS | 9 |
| 2024 | DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the CloudabstractDistributed Deep Learning (DDL), as a paradigm, dictates the use of GPU-based clusters as the optimal infrastructure for training large-scale Deep Neural Networks (DNNs). However, the high cost of such resources makes them inaccessible to many users. Public cloud services, particularly Spot Virtual Machines (VMs), offer a cost-effective alternative, but their unpredictable availability poses a significant challenge to the crucial checkpointing process in DDL. To address this, we introduce DeepVM, a novel solution that recommends cost-effective cluster configurations by intelligently balancing the use of Spot and On-Demand VMs. DeepVM leverages a four-stage process that analyzes instance performance using the FLOPP (FLoating-point Operations Per Price) metric, performs architecture-level analysis with linear programming, and identifies the optimal configuration for the user-specific needs. Extensive simulations and real-world deployments in the AWS environment demonstrate that DeepVM consistently outperforms other policies, reducing training costs and overall makespan. By enabling cost-effective checkpointing with Spot VMs, DeepVM opens up DDL to a wider range of users and facilitates a more efficient training of complex DNNs. Yoochan Kim, Yonghyeon Cho, Awais Khan 0002, Ki-Dong Kang, Baik-Song An, Myung-Hoon Cha, Hong-Yeon Kim, Youngjae Kim 0001 |
CCGrid | 10 |
| 2024 | BandSlim: A Novel Bandwidth and Space-Efficient KV-SSD with an Escape-from-Block ApproachabstractThe Key-Value Solid State Drive (KV-SSD) represents a significant evolution in storage device interfaces by accommodating non-page-aligned key-value pairs, a departure from conventional models. However, KV-SSDs encounter challenges as their specialized data transfer and packing requirements conflict with established storage protocols like NVMe, which are designed around fixed memory page units. This discord leads to inefficient data movement and increased NAND page write I/Os, which in turn escalates network traffic and degrades both performance and NAND efficiency. To tackle these challenges, this paper introduces BandSlim, a novel solution equipped with two methods to streamline bandwidth during I/O transmission: (i) a fine-grained inline value transfer utilizing NVMe commands for bandwidth-efficient value transfer, and (ii) a selective value packing strategy combined with a backfilling policy to reduce NAND page write I/Os. We integrated BandSlim on a state-of-the-art FPGA-based LSM-tree KV-SSD, utilizing the Cosmos+ OpenSSD platform. Our comprehensive evaluations illustrate that BandSlim achieves a remarkable reduction in PCIe traffic of up to 97.9% and NAND page write counts by up to 98.1% compared to the NVMe-based KV-SSD without employing BandSlim. Junhyeok Park 0002, Chang-Gyu Lee, Soon Hwang, Soonyeal Yang, Jungki Noh, Woosuk Chung, Junghee Lee 0004, Youngjae Kim 0001 |
ICPP | 8 |
| 2024 | Towards A Unified Garbage Collection Strategy in ZNS Key-Value Store File Systems Using Same-Victim GCabstractZoned Namespace (ZNS) SSDs are gaining traction for eliminating in-device GC and enabling application-aware data management. BlobDB, an enhanced RocksDB with key-value separation, reduces compaction overhead but suffers from the GC over GC (GoG) problem, causing redundant data copying during BlobDB’s GC and Zone Cleaning (ZC). To address this, this paper proposes Same-Victim GC, aligning the victims and sizes of both GCs. Specifically, we introduce the BlobDB-Aware Zone Allocation (BAZA) algorithm to allocate blob files by creation order, eliminate victim file mismatch between two GCs, and $Z_{-} C u t o f f$ to minimize BlobDB’s GC size without additional overhead. Implemented in ZenFS v2.14 and RocksDB v7.4, our solution eliminates valid data copying, doubles compaction performance, and improves space utilization by $1.28 \times$. Hamin Hwangbo, Joseph Ro, Sungjin Byeon, Safdar Jamil, Jun Young Han, Jooyoung Hwang, Youngjae Kim 0001 |
MASCOTS | 7 |
| 2024 | An Analytical Model-based Capacity Planning Approach for Building CSD-based Storage SystemsabstractThe data movement in large-scale computing facilities (from compute nodes to data nodes) is categorized as one of the major contributors to high cost and energy utilization. To tackle it, in-storage processing (ISP) within storage devices, such as Solid-State Drives (SSDs), has been explored actively. The introduction of computational storage drives (CSDs) enabled ISP within the same form factor as regular SSDs and made it easy to replace SSDs within traditional compute nodes. With CSDs, host systems can offload various operations such as search, filter, and count. However, commercialized CSDs have different hardware resources and performance characteristics. Thus, it requires careful consideration of hardware, performance, and workload characteristics for building a CSD-based storage system within a compute node. Therefore, storage architects are hesitant to build a storage system based on CSDs as there are no tools to determine the benefits of CSD-based compute nodes to meet the performance requirements compared to traditional nodes based on SSDs. In this work, we proposed an analytical model-based storage capacity planner called CsdPlan for system architects to build performance-effective CSD-based compute nodes. Our model takes into account the performance characteristics of the host system, targeted workloads, and hardware and performance characteristics of CSDs to be deployed and provides optimal configuration based on the number of CSDs for a compute node. Furthermore, CsdPlan estimates and reduces the total cost of ownership (TCO) for building a CSD-based compute node. To evaluate the efficacy of CsdPlan , we selected two commercially available CSDs and four representative big data analysis workloads. Hongsu Byun, Safdar Jamil, Jungwook Han, Sungyong Park, Myungcheol Lee, Changsoo Kim, Beongjun Choi, Youngjae Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2023 | KV-CSD: A Hardware-Accelerated Key-Value Store for Data-Intensive ApplicationsabstractPopular software key-value stores such as LevelDB and RocksDB are often tailored for efficient writing. Yet, they tend to also perform well on read operations. This is because while data is initially stored in a format that favors writes, it is later transformed by the DB in the background into a format that better accommodates reads. Write-optimized key-value stores can still block writes. This happens when those background workers cannot keep up with the foreground insertion workload.This paper advocates for a hardware-accelerated key-value store, enabling performance-critical operations, like background data reorganization and queries, to execute directly on storage instead of a host as existing key-value stores do. This better hides background work latency, prevents it from blocking foreground writes, and improves overall I/O efficiency. Our prototype, called KV-CSD, is a key-value based computational storage device consisting of an NVMe SSD and a System-on-a-Chip (SoC) that implements an ordered key-value store atop the SSD. Through offloaded processing, KV-CSD streamlines data insertion, reduces host-device data movement for both background data reorganization and query processing, and shows up to 10.6× lower write times and up to 7.4× faster queries compared to the current state-of-the-art software key-value stores on a real scientific dataset. Inhyuk Park, Qing Zheng, Dominic Manno, Soonyeal Yang, Jason Lee 0004, David Bonnie, Bradley W. Settlemyer, Youngjae Kim 0001, Woosuk Chung, Gary Grider |
CLUSTER | 8 |
| 2023 | A Free-Space Adaptive Runtime Zone-Reset Algorithm for Enhanced ZNS EfficiencyabstractWhile the state-of-the-art runtime zone-reset algorithm of the ZenFS in RocksDB is optimized for the performance of Zone Namespace (ZNS) SSDs, it does not take into account the lifetime constraint of ZNS SSDs. To address this issue, we present FAR, a Free-space Adaptive Runtime Zone-Reset algorithm for ZenFS, which dynamically adjusts the frequency of runtime zone-reset calls based on the available free-space in the ZNS SSD. We developed FAR with the ZenFS of RocksDB using a ZNS SSD prototype based on the Cosmos+ OpenSSD platform and compared it with the state-of-the-art runtime zone-reset algorithm used in ZenFS. Our extensive evaluations demonstrate that FAR improves the lifetime of ZNS SSD by 2x without compromising performance. Sungjin Byeon, Joseph Ro, Safdar Jamil, Jeong-Uk Kang, Youngjae Kim 0001 |
HotStorage | 5 |
| 2023 | OCTOKV: An Agile Network-Based Key-Value Storage System with Robust Load OrchestrationabstractIn this paper, we propose OctoKV, an innovative network-based key-value storage system. OctoKV addresses the repetitive address translation overhead associated with traditional key-value stores running on file systems on the client side. To mitigate this overhead, we implemented the key-value store on the server side using NVMe-oF and a user-level NVMe driver. In particular, we employed fine-grained resource monitoring and load balancing based on heuristics to optimize I/O performance. OctoKV is deployed on a Linux cluster with Intel SPDK. The extensive evaluation shows that OctoKV achieves lower I/O response times in comparison to traditional approaches where key-value stores run on the client side. Also, the proposed load balancing strategies efficiently enhance I/O response times by equally distributing the workload from overloaded cores to other cores. Yeohyeon Park, Junhyeok Park 0002, Awais Khan 0002, Chang-Gyu Lee, Woosuk Chung, Youngjae Kim 0001 |
MASCOTS | 7 |
| 2023 | Iterator Interface Extended LSM-tree-based KVSSD for Range QueriesabstractKey-Value SSD (KVSSD) has shown great potential for several important classes of emerging data stores due to its high throughput and low latency. When designing a key-value store with range queries, an LSM-tree is considered a better choice than a hash table due to its key ordering. However, the design space for range queries in LSM-tree-based KVSSDs has yet to be explored, despite range queries being one of the most demanding features. In this paper, we investigate the design constraints in LSM-tree-based KVSSDs from the perspective of range queries and propose three design principles. Based on these principles, we present IterKVSSD, an Iterator interface extended LSM-tree-based KVSSD for range queries. We implement IterKVSSD on OpenSSD Cosmos+, and our evaluation shows that it increases range query throughput by up to 4.13× and 7.22× for random and sequential key distributions, respectively, compared to existing KVSSDs. Chang-Gyu Lee, Donghyun Min, Inhyuk Park, Woosuk Chung, Anand Sivasubramaniam, Youngjae Kim 0001 |
SYSTOR | 7 |
| 2023 | A Multi-tenant Key-value SSD with Secondary Index for Search Query Processing and AnalysisabstractKey-value SSDs (KVSSDs) introduced so far are limited in their use as an alternative to the key-value store running on the host due to the following technical limitations. First, they were designed only for a single tenant, limiting the use of multiple tenants. Second, they mainly focused on designing indexes for primary key-based searches, without supporting various queries using a combination of primary key and non-primary attribute-based searches. This article proposes Cerberus , a Log Structured Merged (LSM) tree-based KVSSD armed with (1) namespace and performance isolation for multiple tenants in a multi-tenant environment and (2) capability for processing non-primary attribute-based search queries. Specifically, Cerberus identifies the tenant’s namespace and splits a single large LSM-tree into namespace-specific LSM-tree indexes for tenants. Cerberus also manages secondary LSM-tree indexes to enable non-primary attribute-based data access and fast search query processing. With the SSD-internal CPU/DRAM resources, Cerberus supports non-primary attribute-based search queries and handles complex queries that are combined with search and computing operations. We prototyped Cerberus on the Cosmos+ OpenSSD platform. When there are multiple tenants, Cerberus exhibits up to 2.9× higher read throughput and negligible write overhead compared to existing KVSSD. Cerberus also shows lower latency by up to 9.31× for non-primary attribute-based queries. Donghyun Min, Chaewon Moon, Awais Khan 0002, Changhwan Youn, Woosuk Chung, Youngjae Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2022 | Compaction-aware zone allocation for LSM based key-value store on ZNS SSDsabstractUnlike traditional block-based SSDs, Zoned Namespace (ZNS) SSDs expose storage through the zoned block interface, completely eliminating the need for in-device garbage collection (GC) and relinquishing this responsibility to applications. As a result, application-aware data placement decisions give the opportunity for applications on the host to perform efficient GC. Meanwhile, RocksDB for ZNS SSD places data with similar invalidation times (lifetimes) in the same zone through ZenFS (a user-level file system) using the Lifetime-based Zone Allocation algorithm (LIZA), and minimizes the GC overhead of valid data copy when reclaiming a zone. However, LIZA, which allocates zones by predicting the lifetime of each SSTable according to the level of the hierarchical structure of the LSM-tree, is very inefficient in minimizing the write amplification (WA) problem due to inaccurate predictions of SSTable lifetimes. Instead, based on our observation that the deletion time of SSTables in the LSM-tree is solely determined by the compaction process, we propose a novel Compaction-Aware Zone Allocation algorithm (CAZA) that allows the newly created SSTables to be deleted together after merging in the future. CAZA is implemented in RocksDB's ZenFS and our extensive evaluations show that CAZA significantly reduces the WA overhead compared to LIZA. Hee-Rock Lee, Chang-Gyu Lee, Youngjae Kim 0001 |
HotStorage | 4 |
| 2022 | DENOVA: Deduplication Extended NOVA File SystemabstractThis paper shows mathematically and experimentally that inline deduplication is not suitable for file systems on ultra-low latency Intel Optane DC PM devices in terms of performance, and proposes DeNova, an offline deduplication specially designed for log-structured NVM file systems such as NOVA. DeNova offers high-performance and low-latency I/O processing and executes deduplication in the background without interfering with foreground I/Os. DeNova employs DRAM-free persistent deduplication metadata, favoring CPU cache line, and ensures failure consistency on any system failure. We implement DeNova in the NOVA file system. Evaluation with DeNova confirms a negligible performance drop of baseline NOVA of less than 1%, while gaining high storage space savings. Extensive experiments show DeNova is failure consistent in all failure scenario cases. Hyungjoon Kwon, Yonghyeon Cho, Awais Khan 0002, Yeohyeon Park, Youngjae Kim 0001 |
IPDPS | 5 |
| 2022 | A Content-Based Ransomware Detection and Backup Solid-State Drive for Ransomware DefenseabstractRansomware is a growing concern in business and government because it causes immediate financial damages or loss of important data. There is a way to detect and block ransomware in advance, but evolved ransomware can still attack while avoiding detection. Another alternative is to back up the original data. However, existing backup solutions can be under the control of ransomware and backup copies can be destroyed by ransomware. Moreover, backup methods incur storage and performance overhead. In this article, we propose AMOEBA, a device-level backup solution that does not require additional storage for backup. AMOEBA is armed with: 1) a hardware accelerator to run content-based detection algorithms for ransomware detection at high speed and 2) a fine-grained backup control mechanism to minimize space overhead for data backup. For evaluations, we not only implemented AMOEBA using the Microsoft solid-state drive (SSD) simulator but also prototyped it on the OpenSSD-platform. Our extensive evaluations with real ransomware workloads show that AMOEBA has high ransomware detection accuracy with negligible performance overhead. Donghyun Min, Yungwoo Ko, Ryan Walker, Junghee Lee 0004, Youngjae Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Isolating namespace and performance in key-value SSDs for multi-tenant environmentsabstractKey-value SSDs (KVSSDs) implement the storage engine of a key-value store such as log-structured merge-tree (LSM-tree) inside the SSD. However, recent LSM-tree based KVSSDs cannot be used directly in a multi-tenant environment. LSMtree-based KVSSDs are not designed with isolation in mind in terms of namespaces and performance, leading to incorrect data access between concurrent users and poor read performance. In this paper, we propose Iso-KVSSD, a LSM-tree based KVSSD for multi-tenancy by supporting namespace and performance isolation. The Iso-KVSSD performs access control based on the user's namespace and constructs per-namespace dedicated LSM-trees for users. We implement the Iso-KVSSD on Cosmos+ OpenSSD in a Linux environment and evaluate performance with Put() and Get() workloads by varying the number of tenants. Our extensive evaluation results showed that Iso-KVSSD has negligible write performance overhead and an average 2.9 times higher read throughput than a baseline that manages one global shared LSM tree between users. Donghyun Min, Youngjae Kim 0001 |
HotStorage | 2 |
| 2021 | An Analysis of System Balance and Architectural Trends Based on Top500 SupercomputersabstractSupercomputer design is a complex, multi-dimensional optimization process, wherein several subsystems need to be reconciled to meet a desired figure of merit performance for a portfolio of applications and a budget constraint. However, overall, the HPC community has been gravitating towards ever more Flops, at the expense of many other subsystems. To draw attention to overall system balance, in this paper, we analyze balance ratios and architectural trends in the world’s most powerful supercomputers. Specifically, we have collected the performance characteristics of systems between 1993 and 2019 based on the Top500 lists and then analyzed their architectures from diverse system design perspectives. Notably, our analysis studies the performance balance of the machines, across a variety of subsystems such as compute, memory, I/O, interconnect, intra-node connectivity and power. Our analysis reveals that balance ratios of the various subsystems need to be considered carefully alongside the application workload portfolio to provision the subsystem capacity and bandwidth specifications, which can help achieve optimal performance. Awais Khan 0002, Hyogi Sim, Sudharshan S. Vazhkudai, Ali Raza Butt, Youngjae Kim 0001 |
HPC Asia | 5 |
| 2021 | Enabling manycore scalability in F2FS metadata for unlink() operationabstractManycore systems enable massive parallel I/O in a single server due to the number of cores. Among file I/O operations in a file system, C. Lee et al. [1] applied range lock in F2FS for parallel data I/O, and showed scalable performance. However, little research has been done on metadata I/O scalability. Soon Hwang, Chang-Gyu Lee, Youngjae Kim 0001 |
SYSTOR | 3 |
| 2021 | A Probabilistic Machine Learning Approach to Scheduling Parallel Loops With Bayesian OptimizationabstractThis article proposes Bayesian optimization augmented factoring self-scheduling (BO FSS), a new parallel loop scheduling strategy. BO FSS is an automatic tuning variant of the factoring self-scheduling (FSS) algorithm and is based on Bayesian optimization (BO), a black-box optimization algorithm. Its core idea is to automatically tune the internal parameter of FSS by solving an optimization problem using BO. The tuning procedure only requires online execution time measurement of the target loop. In order to apply BO, we model the execution time using two Gaussian process (GP) probabilistic machine learning models. Notably, we propose a locality-aware GP model, which assumes that the temporal locality effect resembles an exponentially decreasing function. By accurately modeling the temporal locality effect, our locality-aware GP model accelerates the convergence of BO. We implemented BO FSS on the GCC implementation of the OpenMP standard and evaluated its performance against other scheduling algorithms. Also, to quantify our method's performance variation on different workloads, or workload-robustness in our terms, we measure the minimax regret. According to the minimax regret, BO FSS shows more consistent performance than other algorithms. Within the considered workloads, BO FSS improves the execution time of FSS by as much as 22% and 5% on average. Kyurai Kim, Youngjae Kim 0001, Sungyong Park |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | DISKSHIELD: A Data Tamper-Resistant Storage for Intel SGXabstractWith the increasing importance of data, the threat of malware which destroys data has been increasing. If malware acquires the highest software privilege, any attempt to detect and remove malware can be disabled. In this paper, we propose DISKSHIELD, a secure storage framework. DISKSHIELD uses Intel SGX to provide Trusted Execution Environment (TEE) to the host, implements the file system into SSD firmware that provides a Trusted Computing Base (TCB), and uses a two-way authentication mechanism to securely transfer data from the host TEE to the SSD TCB against data tampering attacks. This design frees DISKSHIELD from attacks to the kernel. To show the efficacy of DISKSHIELD, we prototyped a DISKSHIELD system by modifying Intel IPFS and developing a device file system on the Jasmine OpenSSD Platform in a Linux environment. Our results show that DISKSHIELD provides strong data tamper resistance the throughput of read and write is on average to 28%, 19% lower than IPFS. Jinwoo Ahn, Junghee Lee 0004, Yungwoo Ko, Donghyun Min, Jiyun Park, Sungyong Park, Youngjae Kim 0001 |
AsiaCCS | 7 |
| 2020 | Position: SGX-SSD: A Policy-based Versioning SSD with Intel SGX
Jinwoo Ahn, Jinhoon Lee, Yungwoo Ko, Donghyun Min, Junghee Lee 0004, Youngjae Kim 0001 |
HotStorage | 7 |
| 2020 | Position: GPUKV: Towards a GPU-Driven Computing on Key-Value SSD
Min-Gyo Jeong, Chang-Gyu Lee, DongGyu Park, Sungyong Park, Youngjae Kim 0001, Jungki Noh, Woosuk Chung, Kyoung Park |
HotStorage | 5 |
| 2020 | A NUMA-aware NVM File System Design for Manycore Server ApplicationsabstractNOVA, a state-of-the-art NVM-based file system, is known to have scalability bottlenecks when multiple I/O threads read/write data simultaneously. Recent studies have identified the cause as the coarse-grained lock adopted by NOVA to provide consistency, and proposed fine-grained range-based locks to improve the scalability of NOVA. However, these variants of NOVA only scale on Uniform Memory Access (UMA) architecture and do not scale on Non-Uniform Memory Access (NUMA) architecture. This is because NOVA has no NUMA-aware memory allocation policy and still uses non-scalable file data structures. In this paper, we propose a NUMA-aware NOVA file system which virtualizes the NVM devices located across NUMA nodes so that they can be used as a single address space. The proposed file system adopts a local-first placement policy where file data and metadata are placed preferentially on the local NVM device to reduce the remote access problem. In addition, the lock-free per-core data structures proposed in this file system allow data to be updated concurrently while mitigating the remote memory access. Extensive evaluations show that our NUMA-aware NOVA for parallel writing is scalable with respect to the increased core count and outperforms vanilla NOVA by 2.56-19.18 times. June-Hyung Kim, Youngjae Kim 0001, Safdar Jamil, Sungyong Park |
MASCOTS | 2 |
| 2020 | Crocus: Enabling Computing Resource Orchestration for Inline Cluster-Wide Deduplication on Scalable Storage SystemsabstractInline deduplication dramatically improves storage space utilization. However, it degrades I/O throughput due to computeintensive deduplication operations such as chunking, fingerprinting or hashing of chunk content, and redundant lookup I/Os over the network in the I/O path. In particular, the fingerprint or hash generation of content contributes largely to the degraded I/O throughput and is computationally expensive. In this article, we propose CROCUS, a framework that enables compute resource orchestration to enhance cluster-wide deduplication performance. In particular, CROCUS takes into account all compute resources such as local and remote {CPU, GPU} by managing decentralized compute pools. An opportunistic Load-Aware Fingerprint Scheduler (LAFS), distributes and offloads compute-intensive deduplication operations in a load-aware fashion to compute pools. CROCUS is highly generic and can be adopted in both inline and offline deduplication with different storage tier configurations. We implemented CROCUS in Ceph scale-out storage system. Our extensive evaluation shows that CROCUS reduces the fingerprinting overhead by 86 percent with 4KB chunk size compared to Ceph with baseline deduplication while maintaining high disk-space savings. Our proposed LAFS scheduler, when tested in different internal and external contention scenarios also showed 54 percent improvement over a fixed or static scheduling approach. Prince Hamandawana, Awais Khan 0002, Chang-Gyu Lee, Sungyong Park, Youngjae Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | An Integrated Indexing and Search Service for Distributed File SystemsabstractData services such as search, discovery, and management in scalable distributed environments have traditionally been decoupled from the underlying file systems, and are often deployed using external databases and indexing services. However, modern data production rates, looming data movement costs, and the lack of metadata, entail revisiting the decoupled file system-data services design philosophy. In this article, we present TagIt, a scalable data management service framework aimed at scientific datasets, which can be integrated into prevalent distributed file system architectures. A key feature of TagIt is a scalable, distributed metadata indexing framework, which facilitates a flexible tagging capability to support data discovery. Furthermore, the tags can also be associated with an active operator, for pre-processing, filtering, or automatic metadata extraction, which we seamlessly offload to file servers in a load-aware fashion. We have integrated TagIt into two popular distributed file systems, i.e., GlusterFS and CephFS. Our evaluation demonstrates that TagIt can expedite data search operation by up to 10× over the extant decoupled approach. Hyogi Sim, Awais Khan 0002, Sudharshan S. Vazhkudai, Seung-Hwan Lim, Ali Raza Butt, Youngjae Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | Towards Robust Data-Driven Parallel Loop Scheduling Using Bayesian OptimizationabstractEfficient parallelization of loops is critical to improving the performance of high-performance computing applications. Many classical parallel loop scheduling algorithms have been developed to increase parallelization efficiency. Recently, workload-aware methods were developed to exploit the structure of workloads. However, both classical and workload-aware scheduling methods lack what we call robustness. That is, most of these scheduling algorithms tend to be unpredictable in terms of performance or have specific workload patterns they favor. This causes application developers to spend additional efforts in finding the best suited algorithm or tune scheduling parameters. This paper proposes Bayesian Optimization augmented Factoring Self-Scheduling (BO FSS), a robust data-driven parallel loop scheduling algorithm. BO FSS is powered by Bayesian Optimization (BO), a machine learning based optimization algorithm. We augment a classical scheduling algorithm, Factoring Self-Scheduling (FSS), into a robust adaptive method that will automatically adapt to a wide range of workloads. To compare the performance and robustness of our method, we have implemented BO FSS and other loop scheduling methods on the OpenMP framework. A regret-based metric called performance regret is also used to quantify robustness. Extensive benchmarking results show that BO FSS performs fairly well in most workload patterns and is also very robust relative to other scheduling methods. BO FSS achieves an average of 4% performance regret. This means that even when BO FSS is not the best performing algorithm on a specific workload, it stays within a 4 percentage points margin of the best performing algorithm. Kyurai Kim, Youngjae Kim 0001, Sungyong Park |
MASCOTS | 2 |
| 2019 | iLSM-SSD: An Intelligent LSM-Tree Based Key-Value SSD for Data AnalyticsabstractSeveral key-value stores such as RocksDB and MongoDB are implemented on the file system using the Log-Structured Merge-Tree (LSM-tree). The LSM-tree involves high compaction overhead. To minimize this overhead, WiscKey, the state-of-the-art LSM-tree, separates key and value, appends the value to the Value Log file, and LSM-tree manages only the key and Value Log offset. This minimizes the compaction overhead by reducing the number of SSTables managed by the LSM-tree. However, WiscKey still has a high I/O stack overhead that must go through the OS file system and block-layer. Therefore, this paper proposes iLSM-SSD that implements WiscKey in SSD and supports near-data processing. iLSM-SSD has the following features: (i) iLSM-SSD implements a key-value separation based LSM-tree in a limited memory space inside the SSD. (ii) The Value Log offset update management overhead incurred during the Value Log cleaning has a significant performance impact on CPU and memory-constrained SSD environments. To minimize this overhead, iLSM-SSD implements Scattered Logging, which reuses invalidated Value Log pages on the Value Log. (iii) iLSM-SSD manages the data layout internally. This enables iLSM-SSD to eliminate the need for file system interactions to obtain the data layout for in-storage processing on traditional block-interface-based SSDs. We prototyped the iLSM-SSD on the Cosmos+ OpenSSD platform in a Linux environment. Extensive evaluations with synthetic benchmarks have shown that the PUT performance of iLSM-SSD is 1.6-4 times higher than that of WiscKey implemented in RocksDB. Chang-Gyu Lee, Hyeongu Kang, DongGyu Park, Sungyong Park, Youngjae Kim 0001, Jungki Noh, Woosuk Chung, Kyoung Park |
MASCOTS | 5 |
| 2019 | Write optimization of log-structured flash file system for parallel I/O on manycore serversabstractIn Manycore server environment, we observe the performance degradation in parallel writes and identify the causes as follows - (i) When multiple threads write to a single file simultaneously, the current POSIX-based F2FS file system does not allow this parallel write even though ranges are distinct where threads are writing. (ii) The high processing time of Fsync at file system layer degrades the I/O throughput as multiple threads call Fsync simultaneously. (iii) The file system periodically checkpoints to recover from system crashes. All incoming I/O requests are blocked while the checkpoint is running, which significantly degrades overall file system performance. To solve these problems, first, we propose file systems to employ a fine-grained file-level Range Lock that allows multiple threads to write on mutually exclusive ranges of files rather than the course-grained inode mutex lock. Second, we propose NVM Node Logging that uses NVM as an extended storage space to store file metadata and file system metadata at high speed during Fsync and checkpoint operations. In particular, the NVM Node Logging consists of (i) a fine-grained inode structure to solve the write amplification problem caused by flushing the file metadata in block units and (ii) a Pin Point NAT (Node Address Table) Update, which can allow flushing only modified NAT entries. We implemented Range Lock and NVM Node Logging for F2FS in Linux kernel 4.14.11. Our extensive evaluation at two different types of servers (single socket 10 cores CPU server, multi-socket 120 cores NUMA CPU server) shows significant write throughput improvements in both real and synthetic workloads. Chang-Gyu Lee, Hyunki Byun, Sunghyun Noh, Hyeongu Kang, Youngjae Kim 0001 |
SYSTOR | 5 |
| 2019 | SciSpace: A scientific collaboration workspace for geo-distributed HPC data centers
Awais Khan 0002, Taeuk Kim, Hyunki Byun, Youngjae Kim 0001 |
Future Gener. Comput. Syst. | 4 |
| 2019 | An Analysis Workflow-Aware Storage System for Multi-Core Active Flash ArraysabstractThe need for novel data analysis is urgent in the face of a data deluge from modern applications. Traditional approaches to data analysis incur significant data movement costs, moving data back and forth between the storage system and the processor. Emerging Active Flash devices enable processing on the flash, where the data already resides. An array of such Active Flash devices allows us to revisit how analysis workflows interact with storage systems. By seamlessly blending together the flash storage and data analysis, we create an analysis workflow-aware storage system, AnalyzeThis. Our guiding principle is that analysis-awareness be deeply ingrained in each and every layer of the storage, elevating data analyses as first-class citizens, and transforming AnalyzeThis into a potent analytics-aware appliance. To evaluate the AnalyzeThis system, we have adopted both emulation and simulation approaches. In particular, we have evaluated AnalyzeThis by implementing the AnalyzeThis storage system on top of the Active Flash Array's emulation platform. We have also implemented an event-driven AnalyzeThis simulator, called AnalyzeThisSim, which allows us to address the limitations of the emulation platform, e.g., performance impact of using multi-core SSDs. The results from our emulation and simulation platforms indicate that AnalyzeThis is a viable approach for expediting workflow execution and minimizing data movement. Hyogi Sim, Geoffroy Vallée, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Devesh Tiwari, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | A Robust Fault-Tolerant and Scalable Cluster-Wide Deduplication for Shared-Nothing Storage SystemsabstractDeduplication has been largely employed in distributed storage systems to improve space efficiency. Traditional deduplication research ignores the design specifications of shared-nothing distributed storage systems such as no central metadata bottleneck, scalability, and storage rebalancing. Further, deduplication introduces transactional changes, which are prone to errors in the event of a system failure, resulting in inconsistencies in data and deduplication metadata. In this paper, we propose a robust, fault-tolerant and scalable cluster-wide deduplication that can eliminate duplicate copies across the cluster. We design a distributed deduplication metadata shard which guarantees performance scalability while preserving the design constraints of shared-nothing storage systems. The placement of chunks and deduplication metadata is made cluster-wide based on the content fingerprint of chunks. To ensure transactional consistency and garbage identification, we employ a flag-based asynchronous consistency mechanism. We implement the proposed deduplication on Ceph. The evaluation shows high disk-space savings with minimal performance degradation as well as high robustness in the event of sudden server failure. Awais Khan 0002, Chang-Gyu Lee, Prince Hamandawana, Sungyong Park, Youngjae Kim 0001 |
MASCOTS | 5 |
| 2017 | AnalyzeThat: A Programmable Shared-Memory System for an Array of Processing-In-Memory DevicesabstractProcessing In Memory (PIM), the concept of integrating processing directly with memory, has been attracting a lot of attention since PIM can assist in overcoming the throughput limitation caused by data movement between CPU and memory. The challenge, however, is that it requires the programmers to have a deep understanding of the PIM architecture to maximize the benefits such as data locality and parallel thread execution on multiple PIM devices. In this study, we present AnalyzeThat, a programmable shared-memory system for parallel data processing with PIM devices. Thematic to AnalyzeThat is a rich PIM-Aware Data Structure (PADS), which is an encapsulation that integrally ties together the data, the analysis tasks and the runtime needed to interface with the PIM device array. The PADS abstraction provides (i) a key-value data container that allows programmers to easily store data on multiple PIMs, (ii) a suite of parallel operations with which users can easily implement data analysis applications, and (iii) a runtime, hidden to programmers, which provides the mechanisms needed to overlay both the data and the tasks on the PIM device array in an intelligent fashion, based on PIM-specific information collected from the hardware. We have developed a PIM emulation framework called AnalyzeThat. Our experimental evaluation with representative data analytics applications suggests that the proposed system can significantly reduce the PIM programming effort without losing its technology benefits. Sangkeun Matt Lee, Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai |
CCGrid | 3 |
| 2017 | Vulnerability Analysis of On-Chip Access-Control Memory
Chintan Chavda, Ethan C. Ahn, Yu-Sheng Chen, Youngjae Kim 0001, Kalidas Ganesh, Junghee Lee 0004 |
HotStorage | 4 |
| 2017 | Understanding object-level memory access patterns across the spectrumabstractMemory accesses limit the performance and scalability of countless applications. Many design and optimization efforts will benefit from an in-depth understanding of memory access behavior, which is not offered by extant access tracing and profiling methods. Chao Wang 0056, Nosayba El-Sayed, Xiaosong Ma, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Wei Xue 0003, Daniel Sánchez 0003 |
SC | 5 |
| 2017 | Tagit: an integrated indexing and search service for file systemsabstractData services such as search, discovery, and management in scalable distributed environments have traditionally been decoupled from the underlying file systems, and are often deployed using external databases and indexing services. However, modern data production rates, looming data movement costs, and the lack of metadata, entail revisiting the decoupled file system-data services design philosophy. Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Geoffroy Vallée, Seung-Hwan Lim, Ali Raza Butt |
SC | 2 |
| 2017 | A quantitative model of application slow-down in multi-resource shared systems
Seung-Hwan Lim, Youngjae Kim 0001 |
Perform. Evaluation | 2 |
| 2017 | LAWC: Optimizing Write Cache Using Layout-Aware I/O Scheduling for All Flash StorageabstractFlash memory-based SSD-RAIDs are swiftly replacing conventional hard disk drives by exhibiting improved performance and stability, especially in I/O-intensive environments. However, the variations in latency and throughput occurring due to uncoordinated internal garbage collection cripples further boosting of performance. In addition, the unwanted variations in each SSD can influence the overall performance of the entire flash storage adversely. This performance bottleneck can be essentially reduced by an internal write cache in the RAID controller designed prudently by considering the crucial device characteristics. The state-of-the-art cache write for the RAID controller fails to incorporate device characteristics of flash memory-based SSDs and mitigates the performance gain. In this paper, we propose a novel cache design namely Layout-Aware Write Cache (LAWC) to overcome the performance barrier inculcated by independent garbage collections. LAWC implements (i) improved I/O scheduling for logically partitioned write caches, (ii) a destage write synchronization mechanism to allow individual write caches to flush write blocks into the SSD array in a coordinated manner, and (iii) a two-level hybrid cache algorithm utilizing small front level cache for the improved write cache efficiency. LAWC shows significant reduction in response time by 82.39 percent on RAID-0 and 68.51 percent on RAID-5 types of SSDs when compared with state-of-theart write cache algorithms. Kalidas Ganesh, Youngjae Kim 0001, Monobrata Debnath, Sungyong Park, Junghee Lee 0004 |
IEEE Trans. Computers | 2 |
| 2017 | Optimizing End-to-End Big Data Transfers over Terabits Network InfrastructureabstractWhile future terabit networks hold the promise of significantly improving big-data motion among geographically distributed data centers, significant challenges must be overcome even on today's 100 gigabit networks to realize end-to-end performance. Multiple bottlenecks exist along the end-to-end path from source to sink, for instance, the data storage infrastructure at both the source and sink and its interplay with the wide-area network are increasingly the bottleneck to achieving high performance. In this paper, we identify the issues that lead to congestion on the path of an end-to-end data transfer in the terabit network environment, and we present a new bulk data movement framework for terabit networks, called LADS. LADS exploits the underlying storage layout at each endpoint to maximize throughput without negatively impacting the performance of shared storage resources for other users. LADS also uses the Common Communication Interface (CCI) in lieu of the sockets interface to benefit from hardware-level zero-copy, and operating system bypass capabilities when available. It can further improve data transfer performance under congestion on the end systems using buffering at the source using flash storage. With our evaluations, we show that LADS can avoid congested storage elements within the shared storage resource, improving input/output bandwidth, and data transfer rates across the high speed networks. We also investigate the performance degradation problems of LADS due to I/O contention on the parallel file system (PFS), when multiple LADS tools share the PFS. We design and evaluate a meta-scheduler to coordinate multiple I/O streams while sharing the PFS, to minimize the I/O contention on the PFS. With our evaluations, we observe that LADS with meta-scheduling can further improve the performance by up to 14 percent relative to LADS without meta-scheduling. Youngjae Kim 0001, Scott Atchley, Geoffroy Vallée, Sankeun Lee 0001, Galen M. Shipman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Design and Analysis of Fault Tolerance Mechanisms for Big Data TransfersabstractIncreased growth of the data and the need to move the data between data centers, demands, an efficient data transfer tool which can, not only transfer the data at higher rates but also handle the faults occurred during the transfer. Absence of fault tolerance mechanisms, would need to retransmit the whole data, in case of any error during the transfer. Hence, fault tolerance is an important aspect of big data transfer tools. In this paper, we have considered LADS data transfer tool, which proved to be superior to the existing data transfer tools with respect to the speed of the transfer. However, absence of fault tolerance implementation in LADS might result in data retransmission and congestion issues upon errors. In this paper, we design and analyze fault tolerance mechanisms which can be used with LADS data transfer tool. We have proposed three different fault tolerance mechanisms, File logging, Transaction logging and Universal logging. Also, we have analyzed the space and performance overhead of these fault tolerance mechanisms on LADS data transfer tool. Preethika Kasu, Youngjae Kim 0001, Sungyong Park, Scott Atchley, Geoffroy Vallée |
CLUSTER | 2 |
| 2016 | Time Optimization Modeling for Big Data Placement and Analysis for Geo-Distributed Data CentersabstractBig data storage and sharing are becoming the major demand of the community. To overcome such issues, virtually unified data facilities are being presented with geodistributed data centers by providing the user with the single unified namespace. These unified data storage facilities lack efficient storage and analysis of data. To address these shortcomings in such unified data facilities, we designed and implemented time optimization model which minimizes the job execution time whereby selecting optimal data center with constraints of storage, computational and network bandwidth among all data centers. Our extensive simulation results show that our model provides the optimal decision that leads to minimal end-to-end data placement and analysis times. Awais Khan 0002, Muhammad Attique 0001, Tae-Sun Chung, Youngjae Kim 0001 |
CLUSTER | 4 |
| 2016 | Minimizing CMT Miss Penalty in Selective Page-Level Address Mapping TableabstractFlash Translation Layer (FTL) performs virtual-to-physical address translations and hides the erase-before-write characteristics of Flash. Pure page mapped FTL, which maintains page-level address mappings, is known as the most efficient FTL. However, its huge SRAM requirement to load the entire mapping table limited adoption of its use. In order to reduce SRAM space utilization while maintaining comparable performance, we can selectively cache page-level address mappings into a small SRAM. However, the performance of this approach is limited by miss ratio of cached mapping table (CMT) on SRAM. In this paper, we propose a replica approach of the page-mapped FTL on flash, called Replica to minimize the performance penalty of CMT miss. Ronnie Mativenga, Joon-Young Paik, Junghee Lee 0004, Tae-Sun Chung, Youngjae Kim 0001 |
CLUSTER | 5 |
| 2016 | Android RMI: a user-level remote method invocation mechanism between Android devices
Hee-Eun Kang, Kihyun Jeong, Kwonyong Lee, Sungyong Park, Youngjae Kim 0001 |
J. Supercomput. | 5 |
| 2015 | LADS: Optimizing Data Transfers Using Layout-Aware Data Scheduling
Youngjae Kim 0001, Scott Atchley, Geoffroy Vallée, Galen M. Shipman |
FAST | 1 |
| 2015 | AnalyzeThis: an analysis workflow-aware storage systemabstractThe need for novel data analysis is urgent in the face of a data deluge from modern applications. Traditional approaches to data analysis incur significant data movement costs, moving data back and forth between the storage system and the processor. Emerging Active Flash devices enable processing on the flash, where the data already resides. An array of such Active Flash devices allows us to revisit how analysis workflows interact with storage systems. By seamlessly blending together the flash storage and data analysis, we create an analysis workflow-aware storage system, AnalyzeThis. Our guiding principle is that analysis-awareness be deeply ingrained in each and every layer of the storage, elevating data analyses as first-class citizens, and transforming AnalyzeThis into a potent analytics-aware appliance. We implement the AnalyzeThis storage system atop an emulation platform of the Active Flash array. Our results indicate that AnalyzeThis is viable, expediting workflow execution and minimizing data movement. Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Devesh Tiwari, Ali Anwar 0001, Ali Raza Butt, Lavanya Ramakrishnan |
SC | 2 |
| 2015 | Understanding I/O workload characteristics of a Peta-scale storage system
Youngjae Kim 0001, Raghul Gunasekaran |
J. Supercomput. | 1 |
| 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File SystemsabstractThe Oak Ridge Leadership Computing Facility (OLCF) has deployed multiple large-scale parallel file systems (PFS) to support its operations. During this process, OLCF acquired significant expertise in large-scale storage system design, file system software development, technology evaluation, benchmarking, procurement, deployment, and operational practices. Based on the lessons learned from each new PFS deployment, OLCF improved its operating procedures, and strategies. This paper provides an account of our experience and lessons learned in acquiring, deploying, and operating large-scale parallel file systems. We believe that these lessons will be useful to the wider HPC community. Sarp Oral, James Simmons, Jason Hill, Dustin Leverman, Feiyi Wang, Matthew Ezell, Ross G. Miller, Douglas Fuller, Raghul Gunasekaran, Youngjae Kim 0001, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai, James H. Rogers, David Dillow, Galen M. Shipman, Arthur S. Bland |
SC | 10 |
| 2014 | Coordinating Garbage Collectionfor Arrays of Solid-State DrivesabstractAlthough solid-state drives (SSDs) offer significant performance improvements over hard disk drives (HDDs) for a number of workloads, they can exhibit substantial variance in request latency and throughput as a result of garbage collection (GC). When GC conflicts with an I/O stream, the stream can make no forward progress until the GC cycle completes. GC cycles are scheduled by logic internal to the SSD based on several factors such as the pattern, frequency, and volume of write requests. When SSDs are used in a RAID with currently available technology, the lack of coordination of the SSD-local GC cycles amplifies this performance variance. We propose a global garbage collection (GGC) mechanism to improve response times and reduce performance variability for a RAID of SSDs. We include a high-level design of SSD-aware RAID controller and GGC-capable SSD devices and algorithms to coordinate the GGC cycles. We develop reactive and proactive GC coordination algorithms and evaluate their I/O performance and block erase counts for various workloads. Our simulations show that GC coordination by a reactive scheme improves average response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write-dominated workloads, response time was improved by 69 percent and performance variability was reduced by 71 percent. We show that a proactive GC coordination algorithm can further improve the I/O response times by up to 9 percent and the performance variability by up to 15 percent. We also observe that it could increase the lifetimes of SSDs with some workloads (e.g., Financial) by reducing the number of block erase counts by up to 79 percent relative to a reactive algorithm for write-dominant enterprise workloads. Youngjae Kim 0001, Junghee Lee 0004, Sarp Oral, David Dillow, Feiyi Wang, Galen M. Shipman |
IEEE Trans. Computers | 1 |
| 2014 | HybridPlan: a capacity planning technique for projecting storage requirements in hybrid storage systems
Youngjae Kim 0001, Bhuvan Urgaonkar, Piotr Berman, Anand Sivasubramaniam |
J. Supercomput. | 1 |
| 2013 | Layout-aware I/O Scheduling for terabits data movementabstractMany science facilities, such as the Department of Energy's Leadership Computing Facilities and experimental facilities including the Spallation Neutron Source, Stanford Linear Accelerator Center, and Advanced Photon Source, produce massive amounts of experimental and simulation data. These data are often shared among the facilities and with collaborating institutions. Moving large datasets over the wide-area network (WAN) is a major problem inhibiting collaboration. Next-generation, terabit-networks will help alleviate the problem, however, the parallel storage systems on the endsystem hosts at these institutions can become a bottleneck for terabit data movement. The parallel storage system (PFS) is shared by simulation systems, experimental systems, analysis and visualization clusters, in addition to wide-area data movers. These competing uses often induce temporary, but significant, I/O load imbalances on the storage system, which impact the performance of all the users. The problem is a serious concern because some resources are more expensive (e.g. super computers) or have time-critical deadlines (e.g. experimental data from a light source), but parallel file systems handle all requests fairly even if some storage servers are under heavy load. This paper investigates the problem of competing workloads accessing the parallel file system and how the performance of wide-area data movement can be improved in these environments. First, we study the I/O load imbalance problems using actual I/O performance data collected from the Spider storage system at the Oak Ridge Leadership Computing Facility. Second, we present I/O optimization solutions with layout-awareness on end-system hosts for bulk data movement. With our evaluation, we show that our I/O optimization techniques can avoid the I/O congested disk groups, improving storage I/O times on parallel storage systems for terabit data movement. Youngjae Kim 0001, Scott Atchley, Geoffroy Vallée, Galen M. Shipman |
IEEE BigData | 1 |
| 2013 | Active flash: towards energy-efficient, in-situ data analytics on extreme-scale machines
Devesh Tiwari, Simona Boboila, Sudharshan S. Vazhkudai, Youngjae Kim 0001, Xiaosong Ma, Peter Desnoyers, Yan Solihin |
FAST | 4 |
| 2013 | A Temporal Locality-Aware Page-Mapped Flash Translation Layer
Youngjae Kim 0001, Bhuvan Urgaonkar |
J. Comput. Sci. Technol. | 1 |
| 2013 | Preemptible I/O Scheduling of Garbage Collection for Solid State DrivesabstractUnlike hard disks, flash devices use out-of-place updates operations and require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in an I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semipreemptible GC (PGC) scheme that allows GC processing to be preempted while pending I/O requests in the queue are serviced. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-PGC scheme with realistic workloads demonstrates both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-PGC scheme. In addition, we explore opportunities of a new NAND flash device that supports suspend/resume commands for read, write, and erase operations for fully PGC (F-PGC). Our experiments with an F-PGC enabled flash device show that request response time can be improved by up to 14.57% compared to semi-PGC. Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Jongman Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | NVMalloc: Exposing an Aggregate SSD Store as a Memory Partition in Extreme-Scale MachinesabstractDRAM is a precious resource in extreme-scale machines and is increasingly becoming scarce, mainly due to the growing number of cores per node. On future multi-petaflop and exaflop machines, the memory pressure is likely to be so severe that we need to rethink our memory usage models. Fortunately, the advent of non-volatile memory (NVM) offers a unique opportunity in this space. Current NVM offerings possess several desirable properties, such as low cost and power efficiency, but suffer from high latency and lifetime issues. We need rich techniques to be able to use them alongside DRAM. In this paper, we propose a novel approach for exploiting NVM as a secondary memory partition so that applications can explicitly allocate and manipulate memory regions therein. More specifically, we propose an NVMalloc library with a suite of services that enables applications to access a distributed NVM storage system. We have devised ways within NVMalloc so that the storage system, built from compute node-local NVM devices, can be accessed in a byte-addressable fashion using the memory mapped I/O interface. Our approach has the potential to re-energize out-of-core computations on large-scale machines by having applications allocate certain variables through NVMalloc, thereby increasing the overall memory capacity available. Our evaluation on a 128-core cluster shows that NVMalloc enables applications to compute problem sizes larger than the physical memory in a cost-effective manner. It can bring more performance/efficiency gain with increased computation time between NVM memory accesses or increased data access locality. In addition, our results suggest that while NVMalloc enables transparent access to NVM-resident variables, the explicit control it provides is crucial to optimize application performance. Chao Wang 0056, Sudharshan S. Vazhkudai, Xiaosong Ma, Youngjae Kim 0001, Christian Engelmann |
IPDPS | 5 |
| 2012 | Active Flash: Out-of-core data analytics on flash storageabstractNext generation science will increasingly come to rely on the ability to perform efficient, on-the-fly analytics of data generated by high-performance computing (HPC) simulations, modeling complex physical phenomena. Scientific computing workflows are stymied by the traditional chaining of simulation and data analysis, creating multiple rounds of redundant reads and writes to the storage system, which grows in cost with the ever-increasing gap between compute and storage speeds in HPC clusters. Recent HPC acquisitions have introduced compute node-local flash storage as a means to alleviate this I/O bottleneck. We propose a novel approach, Active Flash, to expedite data analysis pipelines by migrating to the location of the data, the flash device itself. We argue that Active Flash has the potential to enable true out-of-core data analytics by freeing up both the compute core and the associated main memory. By performing analysis locally, dependence on limited bandwidth to a central storage system is reduced, while allowing this analysis to proceed in parallel with the main application. In addition, offloading work from the host to the more power-efficient controller reduces peak system power usage, which is already in the megawatt range and poses a major barrier to HPC system scalability. We propose an architecture for Active Flash, explore energy and performance trade-offs in moving computation from host to storage, demonstrate the ability of appropriate embedded controllers to perform data analysis and reduction tasks at speeds sufficient for this application, and present a simulation study of Active Flash scheduling policies. These results show the viability of the Active Flash model, and its capability to potentially have a transformative impact on scientific data analysis. Simona Boboila, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Peter Desnoyers, Galen M. Shipman |
MSST | 2 |
| 2012 | D-factor: a quantitative model of application slow-down in multi-resource shared systemsabstractScheduling multiple jobs onto a platform enhances system utilization by sharing resources. The benefits from higher resource utilization include reduced cost to construct, operate, and maintain a system, which often include energy consumption. Maximizing these benefits, while satisfying performance limits, comes at a price -- resource contention among jobs increases job completion time. In this paper, we analyze slow-downs of jobs due to contention for multiple resources in a system; referred to as dilation factor. We observe that multiple-resource contention creates non-linear dilation factors of jobs. From this observation, we establish a general quantitative model for dilation factors of jobs in multi-resource systems. A job is characterized by a vector-valued loading statistics and dilation factors of a job set are given by a quadratic function of their loading vectors. We demonstrate how to systematically characterize a job, maintain the data structure to calculate the dilation factor (loading matrix), and calculate the dilation factor of each job. We validated the accuracy of the model with multiple processes running on a native Linux server, virtualized servers, and with multiple MapReduce workloads co-scheduled in a cluster. Evaluation with measured data shows that the D-factor model has an error margin of less than 16%. We also show that the model can be integrated with an existing on-line scheduler to minimize the makespan of workloads. Seung-Hwan Lim, Jae-Seok Huh, Youngjae Kim 0001, Galen M. Shipman, Chita R. Das |
SIGMETRICS | 3 |
| 2012 | Workload Characterization and Performance Implications of Large-Scale Blog ServersabstractWith the ever-increasing popularity of Social Network Services (SNSs), an understanding of the characteristics of these services and their effects on the behavior of their host servers is critical. However, there has been a lack of research on the workload characterization of servers running SNS applications such as blog services. To fill this void, we empirically characterized real-world Web server logs collected from one of the largest South Korean blog hosting sites for 12 consecutive days. The logs consist of more than 96 million HTTP requests and 4.7TB of network traffic. Our analysis reveals the following: (i) The transfer size of nonmultimedia files and blog articles can be modeled using a truncated Pareto distribution and a log-normal distribution, respectively; (ii) user access for blog articles does not show temporal locality, but is strongly biased towards those posted with image or audio files. We additionally discuss the potential performance improvement through clustering of small files on a blog page into contiguous disk blocks, which benefits from the observed file access patterns. Trace-driven simulations show that, on average, the suggested approach achieves 60.6% better system throughput and reduces the processing time for file access by 30.8% compared to the best performance of the Ext4 filesystem. Myeongjae Jeon, Youngjae Kim 0001, Jeaho Hwang, Joonwon Lee, Euiseong Seo |
ACM Trans. Web | 2 |
| 2011 | Provisioning a Multi-tiered Data Staging Area for Extreme-Scale MachinesabstractMassively parallel scientific applications, running on extreme-scale supercomputers, produce hundreds of terabytes of data per run, driving the need for storage solutions to improve their I/O performance. Traditional parallel file systems (PFS) in high performance computing (HPC) systems are unable to keep up with such high data rates, creating a storage wall. In this work, we present a novel multi-tiered storage architecture comprising hybrid node-local resources to construct a dynamic data staging area for extreme-scale machines. Such a staging ground serves as an impedance matching device between applications and the PFS. Our solution combines diverse resources (e.g., DRAM, SSD) in such a way as to approach the performance of the fastest component technology and the cost of the least expensive one. We have developed an automated provisioning algorithm that aids in meeting the check pointing performance requirement of HPC applications, by using a least-cost storage configuration. We evaluate our approach using both an implementation on a large scale cluster and a simulation driven by six-years worth of Jaguar supercomputer job-logs, and show that our approach, by choosing an appropriate storage configuration, achieves 41.5% cost savings with only negligible impact on performance. Ramya Prabhakar, Sudharshan S. Vazhkudai, Youngjae Kim 0001, Ali Raza Butt, Mahmut T. Kandemir |
ICDCS | 3 |
| 2011 | Enhancing I/O throughput via efficient routing and placement for large-scale parallel file systemsabstractAs storage systems get larger to meet the demands of petascale systems, careful planning must be applied to avoid congestion points and extract the maximum performance. In addition, the large data sets generated by such systems makes it desirable for all compute resources to have common access to this data without needing to copy it to each machine. This paper describes a method of placing I/O close to the storage nodes to minimize contention on Cray's SeaStar2+ network, and extends it to a routed Lustre configuration to gain the same benefits when running against a center-wide file system. Our experiments using half of the resources of Spider - the center-wide file system at the Oak Ridge Leadership Computing Facility - show that I/O write bandwidth can be improved by up to 45% (from 71.9 to 104 GB/s) for a direct-attached configuration and by 137% (47.6 GB/s to 115 GB/s) for a routed configuration. We demonstrated up to 20.7% reduction in run-time for production scientific applications. With the full Spider system, we demonstrated over 240 GB/s of aggregate bandwidth using our techniques. David Dillow, Galen M. Shipman, Sarp Oral, Youngjae Kim 0001 |
IPCCC | 5 |
| 2011 | A semi-preemptive garbage collector for solid state drivesabstractNAND flash memory is a preferred storage media for various platforms ranging from embedded systems to enterprise-scale systems. Flash devices do not have any mechanical moving parts and provide low-latency access. They also require less power compared to rotating media. Unlike hard disks, flash devices use out-of-update operations and they require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semi-preemptive GC scheme that can preempt on-going GC processing and service pending I/O requests in the queue. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-preemptive GC sheme with realistic workloads demonstrate both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-preemptive GC scheme. Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Feiyi Wang, Jongman Kim |
ISPASS | 2 |
| 2011 | HybridStore: A Cost-Efficient, High-Performance Storage System Combining SSDs and HDDsabstractUnlike the use of DRAM for caching or buffering, certain idiosyncrasies of SSDs make their integration into existing systems non-trivial. Flash memory suffers from limits on its reliability, is an order of magnitude more expensive than the HDD, and can sometimes be as slow as the HDD (due to excessive garbage collection (GC) induced by high intensity of random writes). Given these trade-offs between HDDs and SSDs in terms of cost, performance, and lifetime, the current consensus among several storage experts is to view SSDs not as a replacement for HDD but rather as a complementary device within the high performance storage hierarchy. We design and evaluate such a hybrid system called Hybrid Store to provide: (a) Hybrid Plan: improved capacity planning technique to administrators with the overall goal of operating within cost-budgets and (b) HybridDyn: improved performance/lifetime guarantees during episodes of deviations from expected workloads through two novel mechanisms: write-regulation and fragmentation busting. As an illustrative example of HybridStore's efficacy, Hybrid Plan is able to find the most cost-effective storage configuration for a large scale workload of Microsoft Research and suggest one MLC SSD with ten 7.2K RPM HDDs instead of fourteen 7.2K RPM HDDs only. HybridDyn is able to reduce the average response time for an enterprise scale random-write dominant workload by about 71%as compared to a HDD-based system. Youngjae Kim 0001, Bhuvan Urgaonkar, Piotr Berman, Anand Sivasubramaniam |
MASCOTS | 1 |
| 2011 | Harmonia: A globally coordinated garbage collector for arrays of Solid-State DrivesabstractSolid-State Drives (SSDs) offer significant performance improvements over hard disk drives (HDD) on a number of workloads. The frequency of garbage collection (GC) activity is directly correlated with the pattern, frequency, and volume of write requests, and scheduling of GC is controlled by logic internal to the SSD. SSDs can exhibit significant performance degradations when garbage collection (GC) conflicts with an ongoing I/O request stream. When using SSDs in a RAID array, the lack of coordination of the local GC processes amplifies these performance degradations. No RAID controller or SSD available today has the technology to overcome this limitation. This paper presents Harmonia, a Global Garbage Collection (GGC) mechanism to improve response times and reduce performance variability for a RAID array of SSDs. Our proposal includes a high-level design of SSD-aware RAID controller and GGC-capable SSD devices, as well as algorithms to coordinate the global GC cycles. Our simulations show that this design improves response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write dominant workloads response time was improved by 69% while performance variability was reduced by 71%. Youngjae Kim 0001, Sarp Oral, Galen M. Shipman, Junghee Lee 0004, David Dillow, Feiyi Wang |
MSST | 1 |
| 2011 | A comprehensive study of energy efficiency and performance of flash-based SSD
Seon-Yeong Park, Youngjae Kim 0001, Bhuvan Urgaonkar, Joonwon Lee, Euiseong Seo |
J. Syst. Archit. | 2 |
| 2010 | Functional Partitioning to Optimize End-to-End Performance on Many-core ArchitecturesabstractScaling computations on emerging massive-core supercomputers is a daunting task, which coupled with the significantly lagging system I/O capabilities exacerbates applications' end-to-end performance. The I/O bottleneck often negates potential performance benefits of assigning additional compute cores to an application. In this paper, we address this issue via a novel functional partitioning (FP) runtime environment that allocates cores to specific application tasks - checkpointing, de-duplication, and scientific data format transformation - so that the deluge of cores can be brought to bear on the entire gamut of application activities. The focus is on utilizing the extra cores to support HPC application I/O activities and also leverage solid-state disks in this context. For example, our evaluation shows that dedicating 1 core on an oct-core machine for checkpointing and its assist tasks using FP can improve overall execution time of a FLASH benchmark on 80 and 160 cores by 43.95% and 41.34%, respectively. Sudharshan S. Vazhkudai, Ali Raza Butt, Xiaosong Ma, Youngjae Kim 0001, Christian Engelmann, Galen M. Shipman |
SC | 6 |
| 2009 | DFTL: a flash translation layer employing demand-based selective caching of page-level address mappingsabstractRecent technological advances in the development of flash-memory based devices have consolidated their leadership position as the preferred storage media in the embedded systems market and opened new vistas for deployment in enterprise-scale storage systems. Unlike hard disks, flash devices are free from any mechanical moving parts, have no seek or rotational delays and consume lower power. However, the internal idiosyncrasies of flash technology make its performance highly dependent on workload characteristics. The poor performance of random writes has been a cause of major concern, which needs to be addressed to better utilize the potential of flash in enterprise-scale environments. We examine one of the important causes of this poor performance: the design of the Flash Translation Layer (FTL), which performs the virtual-to-physical address translations and hides the erase-before-write characteristics of flash. We propose a complete paradigm shift in the design of the core FTL engine from the existing techniques with our Demand-based Flash Translation Layer (DFTL), which selectively caches page-level address mappings. We develop a flash simulation framework called FlashSim. Our experimental evaluation with realistic enterprise-scale workloads endorses the utility of DFTL in enterprise-scale storage systems by demonstrating: (i) improved performance, (ii) reduced garbage collection overhead and (iii) better overload behavior compared to state-of-the-art FTL schemes. For example, a predominantly random-write dominant I/O trace from an OLTP application running at a large financial institution shows a 78% improvement in average response time (due to a 3-fold reduction in operations of the garbage collector), compared to a state-of-the-art FTL scheme. Even for the well-known read-dominant TPC-H benchmark, for which DFTL introduces additional overheads, we improve system response time by 56%. Youngjae Kim 0001, Bhuvan Urgaonkar |
ASPLOS | 2 |
| 2008 | A CFD-Based Tool for Studying Temperature in Rack-Mounted ServersabstractTemperature-aware computing is becoming more important in design of computer systems as power densities are increasing and the implications of high operating temperatures result in higher failure rates of components and increased demand for cooling capability. Computer architects and system software designers need to understand the thermal consequences of their proposals, and develop techniques to lower operating temperatures to reduce both transient and permanent component failures. Recognizing the need for thermal modeling tools to support those researches, there has been work on modeling temperatures of processors at the micro-architectural level which can be easily understood and employed by computer architects for processor designs. However, there is a dearth of such tools in the academic/research community for undertaking architectural/systems studies beyond a processor - a server box, rack or even a machine room. In this paper we presents a detailed 3-dimensional computational fluid dynamics based thermal modeling tool, called ThermoStat, for rack-mounted server systems. We conduct several experiments with this tool to show how different load conditions affect the thermal profile, and also illustrate how this tool can help design dynamic thermal management techniques. We propose reactive and proactive thermal management for rack mounted server and isothermal workload distribution for rack. Jeonghwan Choi, Youngjae Kim 0001, Anand Sivasubramaniam, Jelena Srebric, Qian Wang 0029, Joonwon Lee |
IEEE Trans. Computers | 2 |
| 2007 | Modeling and Managing Thermal Profiles of Rack-mounted Servers with ThermoStatabstractHigh power densities and the implications of high operating temperatures on the failure rates of components are key driving factors of temperature-aware computing. Computer architects and system software designers need to understand the thermal consequences of their proposals, and develop techniques to lower operating temperatures to reduce both transient and permanent component failures. Tools for understanding temperature ramifications of designs have been mainly restricted to industry for studying packaging and cooling mechanisms, with little access to such toolsets for academic researchers. Developing such tools is an arduous task since it usually requires cross-cutting areas of expertise spanning architecture, systems software, thermodynamics, and cooling systems. Recognizing the need for such tools, there has been work on modeling temperatures of processors at the micro-architectural level which can be easily understood and employed by computer architects for processor designs. However, there is a dearth of such tools in the academic/research community for undertaking architectural/systems studies beyond a processor - a server box, rack or even a machine room. This paper presents a detailed 3-dimensional computational fluid dynamics based thermal modeling tool, called ThermoStat, for rack-mounted server systems. Using this tool, we model a 20 (each with dual Xeon processors) node rack-mounted server system, and validate it with over 30 temperature sensor measurements at different points in the servers/rack. We conduct several experiments with this tool to show how different load conditions affect the thermal profile, and also illustrate how this tool can help design dynamic thermal management techniques Jeonghwan Choi, Youngjae Kim 0001, Anand Sivasubramaniam, Jelena Srebric, Qian Wang 0029, Joonwon Lee |
HPCA | 2 |
| 2006 | Understanding the performance-temperature interactions in disk I/O of server workloadsabstractThis paper describes the first infrastructure for integrated studies of the performance and thermal behavior of storage systems. Using microbenchmarks running on this infrastructure, we first gain insight into how I/O characteristics can affect the temperature of disk drives. We use this analysis to identify the most promising, yet simple, "knobs" for temperature optimization of high speed disks, which can be implemented on existing disks. We then analyze the thermal profiles of real workloads that use such disk drives in their storage systems, pointing out which knobs are most useful for dynamic thermal management when pushing the performance envelope. Youngjae Kim 0001, Sudhanva Gurumurthi, Anand Sivasubramaniam |
HPCA | 1 |