Jim Diehl

dblp:33/421 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-4783-1103ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 since 2021
YearPublicationVenuePosition
2025 Emerald Tiers: Focusing on SSD+MAID Through a Green Lens
abstract
As the volume of retained data continues to increase, it is important to design primary storage systems that efficiently respond to access requests while also providing strong sustainability by reducing carbon emissions. Just over two decades ago, the Massive Arrays of Idle Disks (MAID) architecture was introduced as an energy-efficient alternative to traditional HDD-based always-on storage, employing aggressive spin-down strategies to reduce power consumption. However, high access latencies and hardware limitations led to its decline. In this work, we propose a tiered SSD+MAID storage model that combines the low-latency advantages of SSDs with the energy and carbon efficiency of a MAID system, thus offering a modern alternative to MAID while achieving lower carbon emissions than all-SSD storage. To assess the sustainability impact of such a tiered storage system, we develop a comprehensive carbon emission model that incorporates access patterns, update behaviors, and HDD spin-up dynamics. This model captures both operational and embodied carbon costs, enabling evaluations of primary storage with sustainability in mind. Through real-world workloads, we evaluate the proposed SSD+MAID system and show that it can provide a good trade-off between performance, price, and sustainability.
Zhaokang Ke, Jim Diehl, Ya-Shu Chen, David Hung-Chang Du
HotStorage2
2023 K8sES: Optimizing Kubernetes with Enhanced Storage Service-Level Objectives
abstract
Kubernetes (k8s) is a system for managing containerized applications across multiple hosts. It offers automatic deployment, maintenance, scaling, and resource management for applications. Applications in k8s usually have different storage requirements in the form of service-level objectives (SLOs). However, the current k8s storage management has several limitations which cause explicit performance and cost overhead. K8s administrators have to configure storage in advance manually, and users must know configurations and capabilities of provided storage. Users' storage SLOs can be easily violated in k8s.In this paper, we design and implement k8s Enhanced Storage (k8sES) which efficiently supports applications with various storage SLOs along with all other requirements in the Kubernetes environment. We design and incorporate storage scheduling as part of the node scheduling process in k8s. Applications will be scheduled onto the correct nodes and storage without intervention from either users or administrators. Proper storage resources will be dynamically carved based on users' storage SLOs. In addition, we provide a tool to monitor the I/O activities of both applications and storage devices in k8sES. The evaluation shows that k8sES can better meet users' storage SLOs along with other requirements. Also, k8sES can achieve higher resource utilization efficiency with overhead similar to that of the current k8s.
Hao Wen 0001, Zhichao Cao 0002, Bingzhe Li, David Hung-Chang Du, Ayman Abouelwafa, Doug Voigt, Shiyong Liu, Jim Diehl, Fenggang Wu
ICCD8
2021 TrackLace: Data Management for Interlaced Magnetic Recording
abstract
Interlaced Magnetic Recording (IMR) is a promising technology which achieves higher data density and lower write amplification (WA) than Shingled Magnetic Recording (SMR). In IMR, top tracks and bottom tracks are interlaced so each bottom track is partially overlapped with two adjacent top tracks. Top tracks can be updated without any WA, but bottom track updates require reading and rewriting of affected valid data on the two neighboring top tracks. There are few published studies discussing WA in IMR drives. We propose TrackLace to reduce WA for IMR. TrackLace consists of three techniques: Z-Alloc allocates user data to the tracks in alternating directions and spreads unallocated tracks among allocated tracks; Top-Buffer opportunistically utilizes unallocated top tracks to buffer bottom track updates; and Block-Swap progressively swaps bottom track hot data with top track cold data during high space utilization. To further optimize TrackLace performance, we propose a virtual frame design that can keep the relocated block (due to Top-Buffer or Block-Swap) close to its original location and an adaptive buffering mechanism that can avoid unnecessary redirections depending on the write locality. Evaluations show that TrackLace can reduce WA by 45 percent and lower average latency by 31percent compared with baseline schemes.
Fenggang Wu, Bingzhe Li, Baoquan Zhang, Zhichao Cao 0002, Jim Diehl, Hao Wen 0001, David Hung-Chang Du
IEEE Trans. Computers5
2019 TDDFS: A Tier-Aware Data Deduplication-Based File System
abstract
With the rapid increase in the amount of data produced and the development of new types of storage devices, storage tiering continues to be a popular way to achieve a good tradeoff between performance and cost-effectiveness. In a basic two-tier storage system, a storage tier with higher performance and typically higher cost (the fast tier) is used to store frequently-accessed (active) data while a large amount of less-active data are stored in the lower-performance and low-cost tier (the slow tier). Data are migrated between these two tiers according to their activity. In this article, we propose a Tier-aware Data Deduplication-based File System, called TDDFS, which can operate efficiently on top of a two-tier storage environment. Specifically, to achieve better performance, nearly all file operations are performed in the fast tier. To achieve higher cost-effectiveness, files are migrated from the fast tier to the slow tier if they are no longer active, and this migration is done with data deduplication. The distinctiveness of our design is that it maintains the non-redundant (unique) chunks produced by data deduplication in both tiers if possible. When a file is reloaded (called a reloaded file) from the slow tier to the fast tier, if some data chunks of the file already exist in the fast tier, then the data migration of these chunks from the slow tier can be avoided. Our evaluation shows that TDDFS achieves close to the best overall performance among various file-tiering designs for two-tier storage systems.
Zhichao Cao 0002, Hao Wen 0001, Xiongzi Ge, Jim Diehl, David Hung-Chang Du
ACM Trans. Storage5
2018 Data Management Design for Interlaced Magnetic Recording
Fenggang Wu, Baoquan Zhang, Zhichao Cao 0002, Hao Wen 0001, Bingzhe Li, Jim Diehl, David Hung-Chang Du
HotStorage6
2017 Kinetic Action: Performance Analysis of Integrated Key-Value Storage Devices vs. LevelDB Servers
abstract
With the rise of cloud storage and many data intensive applications, there is an unprecedented growth in the volume of unstructured data. In response, key-value object storage is becoming more popular for the ease with which it can store, manage, and retrieve large amounts of this data. Seagate recently launched Kinetic direct-access-over-Ethernet hard drives which incorporate a LevelDB key-value store inside each drive. In this work, we evaluate these drives using micro as well as macro benchmarks to help understand the performance limits, trade-offs, and implications of replacing traditional hard drives with Kinetic drives in data centers and high performance systems. We perform in-depth throughput and latency benchmarking of these Kinetic drives (each acting as a tiny independent server) from a client machine connected to them via Ethernet. We compare these results to a SATA-based and a faster SAS-based traditional server running LevelDB. Our sample Kinetic drives are CPU-bound, but they still average sequential write throughput of 63 MB/sec and sequential read throughput of 78 MB/sec for 1 MB value sizes. They also demonstrate unique Kinetic features including direct disk-to-disk data transfer. Our macro benchmarking using the Yahoo Cloud Serving Benchmark (YCSB) shows that mid-range LevelDB servers outperform the Kinetic drives for several workloads; however, this is not always the case. For larger value sizes, even these first generation sample Kinetic drives outperform a full server for several different workloads.
Manas Minglani, Jim Diehl, Bingzhe Li, Dongchul Park, David J. Lilja, David Hung-Chang Du
ICPADS2
2016 VNRE: Flexible and Efficient Acceleration for Network Redundancy Elimination
abstract
Network Redundancy Elimination (NRE) aims to improve network performance by identifying and removing repeated transmission of duplicate content from remote servers. Using a Content-Defined Chunking (CDC) policy, an inline NRE process can obtain a higher Redundancy Elimination (RE) ratio but may suffer from a considerably higher computational requirement than fixed-size chunking. Additionally, the existing work on NRE is either based on IP packet level redundancy elimination or rigidly adopting a CDC policy with a static empirically-decided expected chunk size. These approaches make it difficult for conventional NRE MiddleBoxes to achieve both high network throughput to match the increasing line speeds and a high RE ratio at the same time. In this paper we present a design and implementation of an inline NRE appliance which incorporates an improved FPGA-based scheme to speed up CDC processing to match the ever increasing network line speeds while simultaneously obtaining a high RE ratio. The overhead of Rabin fingerprinting, which is a key component of CDC, is greatly reduced through the use of a record table and registers in the FPGA. To efficiently utilize the hardware resources, the whole NRE process is handled by a Virtualized NRE (VNRE) controller. The uniqueness of this VNRE that we developed lies in its ability to exploit the redundancy patterns of different TCP flows and customize the chunking process to achieve a higher RE ratio. VNRE will first decide if the chunking policy should be either fixed-size chunking or CDC. Then VNRE decides the expected chunk size for the corresponding chunking policy based on the TCP flow patterns. Implemented in a partially reconfigurable FPGA card, our trace driven evaluation demonstrates that the chunking throughput for CDC in one FPGA processing unit outperforms chunking running in a virtual CPU by nearly 3X. Moreover, through the differentiation of chunking policies for each flow, the overall throughput of the VNRE appliance outperforms one with static NRE configurations by 6X to 57X while still guaranteeing a high RE ratio.
Xiongzi Ge, Chengtao Lu, Jim Diehl, David Hung-Chang Du
IPDPS4
2007 GreenStor: Application-Aided Energy-Efficient Storage
NagaPramod Mandagere, Jim Diehl, David Hung-Chang Du
MSST2