Feng Chen 0005

dblp:21/3047-5 · DBLP profile ↗
← Back
48ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0002-5641-2536ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 10 first-author · 4 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-authorSoftware engineering, systems software and programming languages · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Encrypted Data Compression with Computational Storage Drives
abstract
Modern data center systems need to achieve several critical goals in security, performance, and cost efficiency. However, realizing these goals simultaneously is highly challenging. In secure data storage systems, a common practice is to first compress and encrypt data on the host side and then transmit it to the storage system using a log-based structure. This approach, unfortunately, leads to increased complexity and performance penalty. As an emerging storage technology, Computational Storage Drives (CSD) can not only offload heavy computation burdens to storage device hardware, but also provide a virtualized logical storage space, creating new optimization opportunities. In this paper, we showcase two unique opportunities enabled by the new CSD technology in data storage management. By replacing ordinary SSDs with CSDs, we can realize efficient one-to-one mapping from host-side blocks to storage-side CSD blocks, eliminating the need for a complex log-based structure and the associated heavy-cost operations, such as garbage collections (GC). Moreover, with a carefully redesigned data format in each compression unit, CSDs can transparently remove redundant data across encrypted snapshots in data-intensive environments, such as databases. We have developed a prototype and conducted experiments on ScaleFlux’s CSD 3000 devices to demonstrate the efficacy of these solutions. We hope that our system investigations in this work provide valuable insight into CSDs and inspire researchers and practitioners to explore additional cases for adopting CSDs to improve the performance and productivity of data center systems.
Linsen Ma, Rui Xie 0006, Feng Chen 0005, Xiaodong Zhang 0001, Tong Zhang 0002
SSDBM3
2026 Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models With Structured Pruning in Resource-Limited Clients
abstract
In this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive collaborative learning, we take the first step to consider the unstructured pruning, varying submodel architectures, knowledge loss, and straggler challenges simultaneously. We propose a novel semiasynchronous collaborative training framework, namely Co-S2P, with data distribution-aware structured pruning and cross-block knowledge transfer mechanism to address the above concerns. Furthermore, we provide theoretical proof that Co-S2P can achieve asymptotic optimal convergence rate of O(1/√ N∗EQ). Finally, we conduct extensive experiments on two types of tasks with a real-world hardware testbed including diverse IoT devices. The experimental results demonstrate that Co-S2P improves accuracy by up to 8.8% and resource utilization by up to 1.2× compared to state-of-the-art methods, while reducing memory consumption by approximately 22% and training time by about 24% on all resource-limited devices.
Xiao Zhang 0015, Feng Chen 0005, Yuan Yuan 0040, Yifei Zou, Mengying Zhao, Jianbo Lu 0001, Dongxiao Yu
IEEE Trans. Mob. Comput.5
2025 HaSiS: A Hardware-assisted Single-index Store for Hybrid Transactional and Analytical Processing
Kecheng Huang, Zhaoyan Shen, Zili Shao, Feng Chen 0005, Tong Zhang 0002
FAST4
2023 Catalyst: Optimizing Cache Management for Large In-memory Key-value Systems
abstract
In-memory key-value cache systems, such as Memcached and Redis, are essential in today's data centers. A key mission of such cache systems is to identify the most valuable data for caching. To achieve this, the current system design keeps track of each key-value item's access and attempts to make accurate estimation on its temporal locality. All it aims is to achieve the highest cache hit ratio. However, as cache capacity quickly increases, the overhead of managing metadata for a massive amount of small key-value items rises to an unbearable level. Put it simply, the current fine-grained, heavy-cost approach cannot continue to scale. In this paper, we have performed an experimental study on the scalability challenge of the current key-value cache system design and quantitatively analyzed the inherent issues related to the metadata operations for cache management. We further propose a key-value cache management scheme, called Catalyst , based on a highly efficient metadata structure, which allows us to make effective caching decisions in a scalable way. By offloading non-essential metadata operations to GPU, we can further dedicate the limited CPU and memory resources to the main service operations for improved throughput and latency. We have developed a prototype based on Memcached. Our experimental results show that our scheme can significantly enhance the scalability and improve the cache system performance by a factor of up to 4.3.
Kefei Wang, Feng Chen 0005
Proc. VLDB Endow.2
2022 Removing Double-Logging with Passive Data Persistence in LSM-tree based Relational Databases
Kecheng Huang, Zhaoyan Shen, Zhiping Jia, Zili Shao, Feng Chen 0005
FAST5
2022 Prism-SSD: A Flexible Storage Interface for SSDs
abstract
The rapid adoption of solid-state drives (SSDs) as a major storage component has been made possible, thanks to their ability to export a standard block I/O interface to the file system and application developers. Meanwhile, this high-level abstraction has been shown to limit the utilization of the devices and the performance of applications running on top of them. Indeed, many optimizations of performance-critical applications bypass the standard block interface and rely on low-level control over SSD internal processes. However, the need to directly manage the physical device significantly increases development complexity and cost, and reduces its portability. Thus, application developers must choose between two extreme options, eithereasy developmentoroptimal performance, without a real possibility to balance between these two objectives. To bridge this gap, we propose aflexible storage interfacethat exports the SSD hardware in three levels of abstraction: 1) as a raw flash media with its low-level details; 2) as a group of functions to manage flash capacity; or 3) as a configurable block device. This multilevel abstraction allows developers to choose the degree in which they desire to control the flash hardware in a manner that best suits the applications’ semantics and performance objectives. We demonstrate the usability of this new model withPrism-SSD—a prototype of this interface as a user-level library on the Open-Channel SSD platform. We use each of the interface’s three abstraction levels to modify the I/O module of three representative applications: 1) a key-value cache system; 2) a user-level file system; and 3) a graph processing engine. Prism-SSD improves application performance by 5%–27%, at varying development costs, between 200 and 3500 lines of code.
Zhaoyan Shen, Feng Chen 0005, Gala Yadgar, Zhiping Jia, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Less is More: De-amplifying I/Os for Key-value Stores with a Log-assisted LSM-tree
abstract
In recent years, Log-Structured Merge Tree (LSMtree) based key-value stores, such as LevelDB and RocksDB, have been widely adopted in data center systems. Though optimized for high-speed write processing, the severe I/O amplification remains a critical constraint that hinders them from reaching their maximum performance potential. Unfortunately, this problem is deeply rooted in the fundamental design of the LSMtree structure. A small number of frequently updated key-value items could quickly pollute the entire tree structure, causing repeated changes in the structure and quickly amplifying the amount of disk IOs across the levels in the tree. In this paper, we present a novel scheme, called Log-assisted LSM-tree (L2SM), to fundamentally address the long-existing I/O amplification problem. L2SM adopts a small-size, multi-level log structure to isolate selected key-value items that have a disruptive effect on the tree structure, accumulates and absorbs the repeated updates in a highly efficient manner, and removes obsolete and deleted key-value items at an early stage. We have prototyped the L2SM structure based on LevelDB. Our evaluation with the YCSB benchmark shows promising results by reducing the amount of disk IOs by up to 40.2%, increasing the throughput by up to 67.4%, and decreasing the average latency by up to 40.1%.
Kecheng Huang, Zhiping Jia, Zhaoyan Shen, Zili Shao, Feng Chen 0005
ICDE5
2021 Understanding Energy Efficiency of Databases on Single Board Computers for Edge Computing
abstract
With the rapid advancement in edge computing, a recent trend is to migrate data processing from data centers to the edge to avoid long data transmission latencies. Databases, as an indispensable component, play a crucial role in efficient data management on edge devices. However, a critical limitation of edge devices is the highly constrained energy resource. Databases often incur a heavy load of CPU and storage I/O activities, which raises a particular concern on power-constrained platforms. In this paper, we have conducted an experimental study on the energy consumption of three representative databases, namely SQLite, LevelDB, and MongoDB, on single board computers for edge computing. We find that by deploying an appropriate database according to specific scenarios, the energy consumption can be reduced by a factor of 58.3, and the bandwidth can be improved by a factor of 54.4. Based on our experimental results, we also present several important system implications associated with our findings. We hope our first-hand data and the obtained insight can provide useful guidance for database and edge system designers and practitioners to develop and deploy energy-efficient databases for edge computing.
Kefei Wang, Feng Chen 0005
MASCOTS3
2021 TSCache: An Efficient Flash-based Caching Scheme for Time-series Data Workloads
abstract
Time-series databases are becoming an indispensable component in today's data centers. In order to manage the rapidly growing time-series data, we need an effective and efficient system solution to handle the huge traffic of time-series data queries. A promising solution is to deploy a high-speed, large-capacity cache system to relieve the burden on the backend time-series databases and accelerate query processing. However, time-series data is drastically different from other traditional data workloads, bringing both challenges and opportunities. In this paper, we present a flash-based cache system design for time-series data, called TSCache . By exploiting the unique properties of time-series data, we have developed a set of optimization schemes, such as a slab-based data management, a two-layered data indexing structure, an adaptive time-aware caching policy, and a low-cost compaction process. We have implemented a prototype based on Twitter's Fatcache. Our experimental results show that TSCache can significantly improve client query performance, effectively increasing the bandwidth by a factor of up to 6.7 and reducing the latency by up to 84.2%.
Kefei Wang, Feng Chen 0005
Proc. VLDB Endow.3
2020 A Study on Nine Years of Bitcoin Transactions: Understanding Real-world Behaviors of Bitcoin Miners and Users
abstract
Bitcoin is the world's first blockchain-based, peer-to-peer cryptocurrency system. Being tremendously successful, the Bitcoin system is designed to support reliable, secure, and trusted transactions between untrusted peers. Since its release in 2009, the Bitcoin system has rapidly grown to an unprecedentedly large scale. However, the real-world behaviors of miners and users in the system and the efficacy of the original Bitcoin system design in the field deployment still remain unclear, hindering us from understanding its internals and developing the next-generation cryptocurrency system. In this paper, we study the behaviors of Bitcoin miners and users and their interactions based on quantitative analysis of more than nine years of Bitcoin transaction history, from its first release on January 3rd, 2009 to April 30th, 2018. We have analyzed over 300 million transaction records to study the transactions' processing, confirmation, and implementation. We have obtained several critical findings regarding how the miners and users exploit the high degree of freedom provided by the Bitcoin system to achieve their own interests. For example, we find that miners often attempt to maximize their profits even by sacrificing system performance; users could try to speed up the transaction processing by mistakenly trading off security for reduced latency. Such unexpected behaviors, to some degree, deviate from the original design purposes of the Bitcoin system and could bring undesirable consequences. Besides revealing several unexpected behaviors of the Bitcoin miners and users in the real world, we have also discussed the associated system implications as well as optimization opportunities in the future.
Binbing Hou, Feng Chen 0005
ICDCS2
2020 Kill Two Birds with One Stone: Auto-tuning RocksDB for High Bandwidth and Low Latency
abstract
Log-Structured Merge (LSM) tree based key-value stores are widely deployed in data centers. Due to its complex internal structures, appropriately configuring a modern key-value data store system, which can have more than 50 parameters with various hardware and system settings, is a highly challenging task. Currently, the industry still heavily relies on a traditional, experience-based, hand-tuning approach for performance tuning. Many simply adopt the default setting out of the box with no changes. Auto-tuning, as a self-adaptive solution, is thus highly appealing for achieving optimal or near-optimal performance in real-world deployment.In this paper, we quantitatively study and compare five optimization methods for auto-tuning the performance of LSM-tree based key-value stores. In order to evaluate the auto-tuning processes, we have conducted an exhaustive set of experiments over RocksDB, a representative LSM-tree data store. We have collected over 12,000 experimental records in 6 months, with about 2,000 software configurations of 6 parameters on different hardware setups. We have compared five representative algorithms, in terms of throughput, the 99th percentile tail latency, convergence time, real-time system throughput, and the iteration process, etc. We find that multi-objective optimization (MOO) methods can achieve a good balance among multiple targets, which satisfies the unique needs of key-value services. The more specific Quality of Service (QoS) requirements users can provide, the better performance these algorithms can achieve. We also find that the number of concurrent threads and the write buffer size are the two most impactful parameters determining the throughput and the 99th percentile tail latency across different hardware and workloads. Finally, we provide system-level explanations for the auto-tuning results and also discuss the associated implications for system designers and practitioners. We hope this work will pave the way towards a practical, high-speed auto-tuning solution for key-value data store systems.
Feng Chen 0005
ICDCS2
2020 From Flash to 3D XPoint: Performance Bottlenecks and Potentials in RocksDB with Storage Evolution
abstract
Storage technologies have undergone continuous innovations in the past decade. The latest technical advancement in this domain is 3D XPoint memory. As a type of Non-volatile Memory (NVM), 3D XPoint memory promises great improvement in performance, density, and endurance over NAND flash memory. Compared to flash based SSDs, 3D XPoint based SSDs, such as Intel's Optane SSD, can deliver unprecedented low latency and high throughput. These properties are particularly appealing to I/O intensive applications. Key-value store is such an important application in data center systems. This paper presents the first, in-depth performance study on the impact of the aforesaid storage hardware evolution to RocksDB, a highly popular key-value store based on Log-structured Merge tree (LSM-tree). We have conducted extensive experiments for quantitative measurements on three types of SSD devices. Besides confirming the performance gain of RocksDB on 3D XPoint SSD, our study also reveals several unexpected bottlenecks in the current key-value store design, which hinder us from fully exploiting the great performance potential of the new storage hardware. Based on our findings, we also present three exemplary case studies to showcase the efficacy of removing these bottlenecks with simple methods, achieving a performance improvement by up to 18.8%. We further discuss the implications of our findings for system designers and users to develop schemes in future optimizations. Our study shows that many of the current LSM-tree based key-value store designs need to be carefully revisited to effectively incorporate the new-generation hardware for realizing high-speed data processing.
Feng Chen 0005
ISPASS2
2020 Put an Elephant into a Fridge: Optimizing Cache Efficiency for In-memory Key-value Stores
abstract
In today's data centers, memory-based key-value systems, such as Memcached and Redis, play an indispensable role in providing high-speed data services. The rapidly growing capacity and quickly falling price of DRAM memory in the past years have enabled us to create a large memory-based key-value store, which is able to serve hundreds of Gigabytes to even Terabytes of key-value data all in memory. Unfortunately, CPU cache in modern processors has not seen a similar growth in capacity, still remaining at the level of a few dozens of Megabytes. Such an extremely low cache-to-memory ratio (less than 0.1%) poses a significant new challenge---the limited CPU cache is becoming a severe performance bottleneck that hinders us from fully exploiting the great potential of high-speed memory-based key-value stores. To address this critical challenge, we propose a highly cache-efficient scheme, called Cavast , to optimize the cache utilization of large-capacity in-memory key-value stores. Our goal is to maximize cache efficiency and system performance without any hardware changes. We first present two light-weight, software-only mechanisms to enable user to indirectly control the cache content at application level. Then we propose a set of optimization policies to address several critical design issues that impair cache's efficacy in the current key-value store systems. By carefully reorganizing the data layout in memory, redesigning the hash indexing structure, and offloading garbage collection, we can effectively improve the utilization of the limited cache space. We have developed a module in Linux as a kernel-level support, and implemented two prototypes based on Memcached and Redis with the proposed Cavast scheme. Our experimental studies show promising results. On a 6-core Intel Xeon processor with only 15-MB cache, we can raise the cache hit ratio up to 82.7% with a very small cache-to-memory ratio (0.023%), and significantly increase the key-value system throughput by a factor of up to 4.2.
Kefei Wang, Feng Chen 0005
Proc. VLDB Endow.3
2020 SlimCache: An Efficient Data Compression Scheme for Flash-based Key-value Caching
abstract
Flash-based key-value caching is becoming popular in data centers for providing high-speed key-value services. These systems adopt slab-based space management on flash and provide a low-cost solution for key-value caching. However, optimizing cache efficiency for flash-based key-value cache systems is highly challenging, due to the huge number of key-value items and the unique technical constraints of flash devices. In this article, we present a dynamic on-line compression scheme, called SlimCache , to improve the cache hit ratio by virtually expanding the usable cache space through data compression. We have investigated the effect of compression granularity to achieve a balance between compression ratio and speed, and we leveraged the unique workload characteristics in key-value systems to efficiently identify and separate hot and cold data. To dynamically adapt to workload changes during runtime, we have designed an adaptive hot/cold area partitioning method based on a cost model. To avoid unnecessary compression, SlimCache also estimates data compressibility to determine whether the data are suitable for compression or not. We have implemented a prototype based on Twitter’s Fatcache. Our experimental results show that SlimCache can accommodate more key-value items in flash by up to 223.4%, effectively increasing throughput and reducing average latency by up to 380.1% and 80.7%, respectively.
Zili Shao, Feng Chen 0005
ACM Trans. Storage3
2019 Reo: Enhancing Reliability and Efficiency of Object-based Flash Caching
abstract
The fast-pace advancement of flash technology has enabled us to build a very large-capacity cache system at a low cost. However, the reliability of flash devices still remains a non-negligible concern, especially for flash-based cache. This is for two reasons. First, corruption of dirty data in cache would cause a permanent loss of user data. Second, warming up a huge-capacity cache would take an excessively long period of time. In this paper, we present a highly reliable, efficient, object-based flash cache, called Reo. Reo is designed to leverage the highly expressible object interface to exploit the rich semantic knowledge of the flash cache manager. Reo has two key mechanisms, differentiated data redundancy and differentiated data recovery, to make a flash cache highly reliable, and in the meantime, still remains space efficient for high cache hit ratio. We have prototyped Reo based on open-osd, an open-source implementation of T10 Object Storage Device (OSD) in Linux. Our experimental results show that compared to uniform data protection, Reo achieves graceful performance degradation and prioritized recovery upon device failures. Compared to full replication, Reo is more space efficient and delivers up to 3.1 times of the cache hit ratio and up to 3.6 times of the bandwidth.
Kefei Wang, Feng Chen 0005
ICDCS3
2019 One Size Never Fits All: A Flexible Storage Interface for SSDs
abstract
The rapid adoption of solid-state drives (SSDs) as a major storage component has been made possible thanks to their ability to export a standard block I/O interface to file system and application developers. Meanwhile, this high-level abstraction has been shown to limit the utilization of the devices and the performance of applications running on top of them. Indeed, many optimizations of performance-critical applications bypass the standard block interface and rely on low-level control over SSD internal processes. However, the need to directly manage the physical device significantly increases development complexity and cost, and reduces its portability. Thus, application developers must choose between two extreme options, either easy development or optimal performance, without a real possibility to balance between these two objectives. To bridge this gap, we propose a flexible storage interface that exports the SSD hardware in three levels of abstraction: as a raw flash media with its low-level details, as a group of functions to manage flash capacity, or as a configurable block device. This multi-level abstraction allows developers to choose the degree in which they desire to control the flash hardware in a manner that best suits their applications' semantics and performance objectives. We demonstrate the usability of this new model with Prism-SSD-a prototype of this interface as a user-level library on the Open-Channel SSD platform. We use each of the interface's three abstraction levels to modify the I/O module of three representative applications: a key-value cache system, a user-level file system, and a graph processing engine. Prism-SSD improves application performance by 5% to 27%, at varying development costs, between 200 and 3,500 lines of code.
Zhaoyan Shen, Feng Chen 0005, Gala Yadgar, Zili Shao
ICDCS2
2019 When NVMe over Fabrics Meets Arm: Performance and Implications
abstract
A growing technology trend in the industry is to deploy highly capable and power-efficient storage servers based on the Arm architecture. An important driving force behind this is storage disaggregation, which separates compute and storage to different servers, enabling independent resource allocation and optimized hardware utilization. The recently released remote storage protocol specification, NVMe-over-Fabrics (NVMeoF), makes flash disaggregation possible by reducing the remote access overhead to the minimum. It is highly appealing to integrate the two promising technologies together to build an efficient Arm based storage server with NVMeoF. In this work, we have conducted a set of comprehensive experiments to understand the performance behaviors of NVMeoF on Arm-based Data Center SoC and to gain insight into the implications of their design and deployment in data centers. Our experiments show that NVMeoF delivers the promised ultra-low latency. With appropriate optimizations on both hardware and software, NVMeoF can achieve even better performance than direct attached storage. Specifically, with appropriate NIC optimizations, we have observed a throughput increase by up to 42.5% and a decrease of the 95th percentile tail latency by up to 14.6%. Based on our measurement results, we also discuss several system implications for integrating NVMeoF on Arm based platforms. Our studies show that this system solution can well balance the computation, network, and storage resources for data-center storage services. Our findings have also been reported to Arm and Broadcom for future optimizations.
Eric Anger, Feng Chen 0005
MSST3
2018 Cascade Mapping: Optimizing Memory Efficiency for Flash-based Key-value Caching
abstract
Flash-based key-value caching plays an important role in Internet services. Compared to in-memory key-value caches, flash-based key-value caches can provide a 10 to 100 times larger cache space, allowing to accommodate more data for a higher hit ratio. However, the current design relies on a simple hash-based indexing structure, which maintains the entire mapping table in DRAM memory. As the cache capacity continues to grow, such an "all-in-memory" approach raises concerns on cost, power, and scalability.
Kefei Wang, Feng Chen 0005
SoCC2
2018 Pacaca: Mining Object Correlations and Parallelism for Enhancing User Experience with Cloud Storage
abstract
Object-based cloud storage presents an unconventional storage model. Exploiting its unique characteristics, such as the strong semantic correlations among objects and the high I/O parallelism potential, can greatly enhance user experience. Unfortunately, current storage optimization techniques, such as the caching and prefetching schemes, are designed for conventional storage and thus are sub-optimal for cloud storage services. In this paper, we propose a client-side cache management framework, called Pacaca, which integrates object clustering, parallelized prefetching, and cost-aware caching to exploit I/O parallelism and object correlations on cloud storage. We first develop an efficient mining scheme, called Frequent Cluster Mining (FCM), to discover object correlations from the access sequence, and then build a prefetching scheme to fetch the correlated objects in parallel. These two schemes are closely coordinated for achieving high prefetching accuracy, proper control on parallelism degree, and effective mis-prefetching detection and handling. After studying the impact of parallelized prefetching on cache management, we further present a cost-aware caching scheme to differentiate low-cost and high-cost objects for efficient caching by leveraging the awareness of parallelism and object correlations. Our experimental results show that our optimization schemes can effectively reduce the access latency, outperforming traditional schemes by up to 58%.
Binbing Hou, Feng Chen 0005
MASCOTS2
2018 SlimCache: Exploiting Data Compression Opportunities in Flash-Based Key-Value Caching
abstract
Flash-based key-value caching is becoming popular in data centers for providing high-speed key-value services. These systems adopt slab-based space management on flash and provide a low-cost solution for key-value caching. However, optimizing cache efficiency for flash-based key-value cache systems is highly challenging, due to the huge number of key-value items and the unique technical constraints of flash devices. In this paper, we present a dynamic on-line compression scheme, called SlimCache, to improve the cache hit ratio by virtually expanding the usable cache space through data compression. We have investigated the effect of compression granularity to achieve a balance between compression ratio and speed, and leveraged the unique workload characteristics in key-value systems to efficiently identify and separate hot and cold data. In order to dynamically adapt to workload changes during runtime, we have designed an adaptive hot/cold area partitioning method based on a cost model. Inorder to avoid unnecessary compression, SlimCache also estimates data compressibility to determine whether the data are suitable for compression or not. We have implemented a prototype based on Twitter's Fatcache. Our experimental results show that SlimCache can accommodate more key-value items in flash by up to 125.9%, effectively increasing throughput and reducing average latency by up to 255.6% and 78.9%, respectively.
Zili Shao, Feng Chen 0005
MASCOTS3
2018 DIDACache: An Integration of Device and Application for Flash-based Key-value Caching
abstract
Key-value caching is crucial to today’s low-latency Internet services. Conventional key-value cache systems, such as Memcached, heavily rely on expensive DRAM memory. To lower Total Cost of Ownership, the industry recently is moving toward more cost-efficient flash-based solutions, such as Facebook’s McDipper [14] and Twitter’s Fatcache [56]. These cache systems typically take commercial SSDs and adopt a Memcached-like scheme to store and manage key-value cache data in flash. Such a practice, though simple, is inefficient due to the huge semantic gap between the key-value cache manager and the underlying flash devices. In this article, we advocate to reconsider the cache system design and directly open device-level details of the underlying flash storage for key-value caching. We propose an enhanced flash-aware key-value cache manager, which consists of a novel unified address mapping module, an integrated garbage collection policy, a dynamic over-provisioning space management, and a customized wear-leveling policy, to directly drive the flash management. A thin intermediate library layer provides a slab-based abstraction of low-level flash memory space and an API interface for directly and easily operating flash devices. A special flash memory SSD hardware that exposes flash physical details is adopted to store key-value items. This co-design approach bridges the semantic gap and well connects the two layers together, which allows us to leverage both the domain knowledge of key-value caches and the unique device properties. In this way, we can maximize the efficiency of key-value caching on flash devices while minimizing its weakness. We implemented a prototype, called DIDACache, based on the Open-Channel SSD platform. Our experiments on real hardware show that we can significantly increase the throughput by 35.5%, reduce the latency by 23.6%, and remove unnecessary erase operations by 28%.
Zhaoyan Shen, Feng Chen 0005, Zili Shao
ACM Trans. Storage2
2018 A Low-cost Disk Solution Enabling LSM-tree to Achieve High Performance for Mixed Read/Write Workloads
abstract
LSM-tree has been widely used in data management production systems for write-intensive workloads. However, as read and write workloads co-exist under LSM-tree, data accesses can experience long latency and low throughput due to the interferences to buffer caching from the compaction, a major and frequent operation in LSM-tree. After a compaction, the existing data blocks are reorganized and written to other locations on disks. As a result, the related data blocks that have been loaded in the buffer cache are invalidated since their referencing addresses are changed, causing serious performance degradations. To re-enable high-speed buffer caching during intensive writes, we propose Log-Structured buffered-Merge tree (simplified as LSbM-tree) by adding a compaction buffer on disks to minimize the cache invalidations on buffer cache caused by compactions. The compaction buffer efficiently and adaptively maintains the frequently visited datasets. In LSbM, strong locality objects can be effectively kept in the buffer cache with minimum or no harmful invalidations. With the help of a small on-disk compaction buffer, LSbM achieves a high query performance by enabling effective buffer caching, while retaining all the merits of LSM-tree for write-intensive data processing and providing high bandwidth of disks for range queries. We have implemented LSbM based on LevelDB. We show that with a standard buffer cache and a hard disk, LSbM can achieve 2x performance improvement over LevelDB. We have also compared LSbM with other existing solutions to show its strong cache effectiveness.
Dejun Teng, Lei Guo 0004, Rubao Lee, Feng Chen 0005, Yanfeng Zhang 0001, Xiaodong Zhang 0001
ACM Trans. Storage4
2017 DIDACache: A Deep Integration of Device and Application for Flash Based Key-Value Caching
Zhaoyan Shen, Feng Chen 0005, Zili Shao
FAST2
2017 LSbM-tree: Re-Enabling Buffer Caching in Data Management for Mixed Reads and Writes
abstract
LSM-tree has been widely used in data management production systems for write-intensive workloads. However, as read and write workloads co-exist under LSM-tree, data accesses can experience long latency and low throughput due to the interferences to buffer caching from the compaction, a major and frequent operation in LSM-tree. After a compaction, the existing data blocks are reorganized and written to other locations on disks. As a result, the related data blocks that have been loaded in the buffer cache are invalidated since their referencing addresses are changed, causing serious performance degradations. In order to re-enable high-speed buffer caching during intensive writes, we propose Log-Structured buffered-Merge tree (simplified as LSbM-tree) by adding a compaction buffer on disks, to minimize the cache invalidations on buffer cache caused by compactions. The compaction buffer efficiently and adaptively maintains the frequently visited data sets. In LSbM, strong locality objects can be effectively kept in the buffer cache with minimum or without harmful invalidations. With the help of a small on-disk compaction buffer, LSbM achieves a high query performance by enabling effective buffer caching, while retaining all the merits of LSM-tree for write-intensive data processing, and providing high bandwidth of disks for range queries. We have implemented LSbM based on LevelDB. We show that with a standard buffer cache and a hard disk, LSbM can achieve 2x performance improvement over LevelDB. We have also compared LSbM with other existing solutions to show its strong effectiveness.
Dejun Teng, Lei Guo 0004, Rubao Lee, Feng Chen 0005, Yanfeng Zhang 0001, Xiaodong Zhang 0001
ICDCS4
2017 Software Support Inside and Outside Solid-State Devices for High Performance and High Efficiency
abstract
In the past decade, flash memory has been in the spotlight across a variety of research communities from circuits to computer systems, and significant progress has been accomplished. This has enabled flash memory to become increasingly pervasive across the entire information technology infrastructure, from consumer electronics to cloud and supercomputing. This paper aims to provide a comprehensive survey on the important advancements and milestones in the domains across flash translation layer (FTL), operating systems, and applications. As the storage device hardware has been quickly commoditized, software becomes increasingly important to tap the potential of flash memory to its full extent. Therefore, a comprehensive survey with a focus on software aspects will be very valuable to the research community and industry. It is our hope that this survey paper will serve as a good reference for system practitioners and researchers.
Feng Chen 0005, Tong Zhang 0002, Xiaodong Zhang 0001
Proc. IEEE1
2017 GDS-LC: A Latency- and Cost-Aware Client Caching Scheme for Cloud Storage
abstract
Successfully integrating cloud storage as a primary storage layer in the I/O stack is highly challenging. This is essentially due to two inherent critical issues: the high and variant cloud I/O latency and the per-I/O pricing model of cloud storage. To minimize the associated latency and monetary cost with cloud I/Os, caching is a crucial technology, as it directly influences how frequently the client has to communicate with the cloud. Unfortunately, current cloud caching schemes are mostly designed to optimize miss reduction as the sole objective and only focus on improving system performance while ignoring the fact that various cache misses could have completely distinct effects in terms of latency and monetary cost. In this article, we present a cost-aware caching scheme, called GDS-LC , which is highly optimized for cloud storage caching. Different from traditional caching schemes that merely focus on improving cache hit ratios and the classic cost-aware schemes that can only achieve a single optimization target, GDS-LC offers a comprehensive cache design by considering not only the access locality but also the object size, associated latency, and price, aiming at enhancing the user experience with cloud storage from two aspects: access latency and monetary cost. To achieve this, GDS-LC virtually partitions the cache space into two regions: a high-priority latency-aware region and a low-priority price-aware region. Each region is managed by a cost-aware caching scheme, which is based on GreedyDual-Size (GDS) and designed for a cloud storage scenario by adopting clean-dirty differentiation and latency normalization. The GDS-LC framework is highly flexible, and we present a further enhanced algorithm, called GDS-LCF , by incorporating access frequency in caching decisions. We have built a prototype to emulate a typical cloud client cache and evaluate GDS-LC and GDS-LCF with Amazon Simple Storage Services (S3) in three different scenarios: local cloud, Internet cloud, and heterogeneous cloud. Our experimental results show that our caching schemes can effectively achieve both optimization goals: low access latency and low monetary cost. It is our hope that this work can inspire the community to reconsider the cache design in the cloud environment, especially for the purpose of integrating cloud storage into the current storage stack as a primary layer.
Binbing Hou, Feng Chen 0005
ACM Trans. Storage2
2017 Understanding I/O Performance Behaviors of Cloud Storage from a Client's Perspective
abstract
Cloud storage has gained increasing popularity in the past few years. In cloud storage, data is stored in the service provider’s data centers, and users access data via the network. For such a new storage model, our prior wisdom about conventional storage may not remain valid nor applicable to the emerging cloud storage. In this article, we present a comprehensive study to gain insight into the unique characteristics of cloud storage and optimize user experiences with cloud storage from a client’s perspective. Unlike prior measurement work that mostly aims to characterize cloud storage providers or specific client applications, we focus on analyzing the effects of various client-side factors on the user-experienced performance. Through extensive experiments and quantitative analysis, we have obtained several important findings. For example, we find that (1) a proper combination of parallelism and request size can achieve optimized bandwidths, (2) a client’s capabilities and geographical location play an important role in determining the end-to-end user-perceivable performance, and (3) the interference among mixed cloud storage requests may cause performance degradation. Based on our findings, we showcase a sampling- and inference-based method to determine a proper combination for different optimization goals. We further present a set of case studies on client-side chunking and parallelization for typical cloud-based applications. Our studies show that specific attention should be paid to fully exploiting the capabilities of clients and the great potential of cloud storage services.
Binbing Hou, Feng Chen 0005, Zhonghong Ou, Ren Wang 0001, Michael P. Mesnier
ACM Trans. Storage2
2016 Optimizing Flash-based Key-value Cache Systems
Zhaoyan Shen, Feng Chen 0005, Zili Shao
HotStorage2
2016 Understanding storage I/O behaviors of mobile applications
abstract
In the past few years, mobile devices quickly gained high popularity in our daily life. Designed for ultra-mobility, these small yet powerful devices are fundamentally distinct from traditional computer systems (e.g., PCs and servers) - from the internal hardware architecture and software stack, to application behaviors. Storage, the slowest component in the I/O stack, plays an important role in mobile systems and can greatly affect user experience. In this paper, we present a set of comprehensive experimental studies on mobile storage and attempt to gain insight on the unique behaviors of mobile applications and characterize the performance properties of underlying mobile storage. In our experiments, we carefully selected 13 representative mobile workloads from 5 different categories. Our studies reveal several unexpected observations on mobile storage. Based on these findings, we further discuss the associated implications to mobile systems and application designers. We hope this work can inspire system architects, application designers, and practitioners to pay specific attention to the high-latency I/O operations, rather than completely relying on the default APIs. We also suggest a further look to new opportunities, such as adopting a faster medium in the mobile system architecture, for future research.
Jace Courville, Feng Chen 0005
MSST2
2016 Understanding I/O performance behaviors of cloud storage from a client's perspective
abstract
Cloud storage has gained increasing popularity in the past few years. In cloud storage, data is stored in the service provider's data centers, and users access data via the network. For such a new storage model, our prior wisdom about conventional storage may not remain valid nor applicable to the emerging cloud storage. In this paper, we present a comprehensive study and attempt to gain insight into the unique characteristics of cloud storage, primarily from the client's perspective. Through extensive experiments and quantitative analysis, we have acquired several interesting, and in some cases unexpected, findings. (1) Parallelizing I/Os and increasing request sizes are keys to improving the performance, but optimal bandwidth may only be achieved with a proper combination of parallelism and request size. (2) Client capabilities, including CPU, memory, and storage, play an unexpectedly important role in determining the achievable performance. (3) A geographically long distance affects client-perceived performance but does not always result in lower bandwidth and longer latency. Based on our experimental studies, we further present a case study on appropriate chunking and parallelization in a cloud storage client. Our studies show that specific attention should be paid to fully exploiting the capabilities of clients and the great potential of cloud storage services.
Binbing Hou, Feng Chen 0005, Zhonghong Ou, Ren Wang 0001, Michael P. Mesnier
MSST2
2016 Internal Parallelism of Flash Memory-Based Solid-State Drives
abstract
A unique merit of a solid-state drive (SSD) is its internal parallelism . In this article, we present a set of comprehensive studies on understanding and exploiting internal parallelism of SSDs. Through extensive experiments and thorough analysis, we show that exploiting internal parallelism of SSDs can not only substantially improve input/output (I/O) performance but also may lead to some surprising side effects and dynamics. For example, we find that with parallel I/Os, SSD performance is no longer highly sensitive to access patterns (random or sequential), but rather to other factors, such as data access interferences and physical data layout. Many of our prior understandings about SSDs also need to be reconsidered. For example, we find that with parallel I/Os, write performance could outperform reads and is largely independent of access patterns, which is opposite to our long-existing common understanding about slow random writes on SSDs. We have also observed a strong interference between concurrent reads and writes as well as the impact of physical data layout to parallel I/O performance. Based on these findings, we present a set of case studies in database management systems, a typical data-intensive application. Our case studies show that exploiting internal parallelism is not only the key to enhancing application performance, and more importantly, it also fundamentally changes the equation for optimizing applications. This calls for a careful reconsideration of various aspects in application and system designs. Furthermore, we give a set of experimental studies on new-generation SSDs and the interaction between internal and external parallelism in an SSD-based Redundant Array of Independent Disks (RAID) storage. With these critical findings, we finally make a set of recommendations to system architects and application designers for effectively exploiting internal parallelism.
Feng Chen 0005, Binbing Hou, Rubao Lee
ACM Trans. Storage1
2015 Hetero-DB: Next Generation High-Performance Database Systems by Best Utilizing Heterogeneous Computing and Storage Resources
Kai Zhang 0006, Feng Chen 0005, Xiaoning Ding, Yin Huai, Rubao Lee, Kaibo Wang, Yuan Yuan 0014, Xiaodong Zhang 0001
J. Comput. Sci. Technol.2
2014 A protected block device for Persistent Memory
abstract
Persistent Memory (PM) technologies, such as Phase Change Memory, STT-RAM, and memristors, are receiving increasingly high interest in academia and industry. PM provides many attractive features, such as DRAM-like speed and storage-like persistence. Yet, because it draws a blurry line between memory and storage, neither a memory- or storage-based model is a natural fit. Best integrating PM into existing systems has become challenging and is now a top priority for many. In this paper we share our initial approach to integrating PM into computer systems, with minimal impact to the core operating system. By adopting a hybrid storage model, all of our changes are confined to a block storage driver, called PMBD, which directly accesses PM attached to the memory bus and exposes a logical block I/O interface to users. We explore the design space by examining a variety of options to achieve performance, protection from stray writes, ordered persistence, and compatibility for legacy file systems and applications. All told, we find that by using a combination of existing OS mechanisms (per-core page table mappings, non-temporal store instructions, memory fences, and I/O barriers), we are able to achieve each of these goals with small performance overhead for both micro-benchmarks and real world applications (e.g., file server and database workloads). Our experience suggests that determining the right combination of existing platform and OS mechanisms is a non-trivial exercise. In this paper, we share both our failed and successful attempts. The final solution that we propose represents an evolution of our initial approach. We have also open-sourced our software prototype with all attempted design options to encourage further research in this area.
Feng Chen 0005, Michael P. Mesnier, Scott Hahn
MSST1
2014 Client-aware cloud storage
abstract
Cloud storage is receiving high interest in both academia and industry. As a new storage model, it provides many attractive features, such as high availability, resilience, and cost efficiency. Yet, cloud storage also brings many new challenges. In particular, it widens the already-significant semantic gap between applications, which generate data, and storage systems, which manage data. This widening semantic gap makes end-to-end differentiated services extremely difficult. In this paper, we present a client-aware cloud storage framework, which allows semantic information to flow from clients, across multiple intermediate layers, to the cloud storage system. In turn, the storage system can differentiate various data classes and enforce predefined policies. We showcase the effectiveness of enabling such client awareness by using Intel's Differentiated Storage Services (DSS) to enhance persistent disk caching and to control I/O traffic to different storage devices. We find that we can significantly outperform LRU-style caching, improving upload bandwidth by 5x and download bandwidth by 1.6x. Further, we can achieve 85% of the performance of a full-SSD solution at only a fraction (14%) of the cost.
Feng Chen 0005, Michael P. Mesnier, Scott Hahn
MSST1
2012 hStorage-DB: Heterogeneity-aware Data Management to Exploit the Full Capability of Hybrid Storage Systems
abstract
As storage systems become increasingly heterogeneous and complex, it adds burdens on DBAs, causing suboptimal performance even after a lot of human efforts have been made. In addition, existing monitoring-based storage management by access pattern detections has difficulties to handle workloads that are highly dynamic and concurrent. To achieve high performance by best utilizing heterogeneous storage devices, we have designed and implemented a heterogeneity-aware software framework for DBMS storage management called hStorage-DB, where semantic information that is critical for storage I/O is identified and passed to the storage manager. According to the collected semantic information, requests are classified into different types. Each type is assigned a proper QoS policy supported by the underlying storage system, so that every request will be served with a suitable storage device. With hStorage-DB, we can well utilize semantic information that cannot be detected through data access monitoring but is particularly important for a hybrid storage system. To show the effectiveness of hStorage-DB, we have implemented a system prototype that consists of an I/O request classification enabled DBMS, and a hybrid storage system that is organized into a two-level caching hierarchy. Our performance evaluation shows that hStorage-DB can automatically make proper decisions for data allocation in different storage devices and make substantial performance improvements in a cost-efficient way.
Rubao Lee, Michael P. Mesnier, Feng Chen 0005, Xiaodong Zhang 0001
Proc. VLDB Endow.4
2011 CAFTL: A Content-Aware Flash Translation Layer Enhancing the Lifespan of Flash Memory based Solid State Drives
Feng Chen 0005, Xiaodong Zhang 0001
FAST1
2011 Essential roles of exploiting internal parallelism of flash memory based solid state drives in high-speed data processing
abstract
Flash memory based solid state drives (SSDs) have shown a great potential to change storage infrastructure fundamentally through their high performance and low power. Most recent studies have mainly focused on addressing the technical limitations caused by special requirements for writes in flash memory. However, a unique merit of an SSD is its rich internal parallelism, which allows us to offset for the most part of the performance loss related to technical limitations by significantly increasing data processing throughput. In this work we present a comprehensive study of essential roles of internal parallelism of SSDs in high-speed data processing. Besides substantially improving I/O bandwidth (e.g. 7.2×), we show that by exploiting internal parallelism, SSD performance is no longer highly sensitive to access patterns, but rather to other factors, such as data access interferences and physical data layout. Specifically, through extensive experiments and thorough analysis, we obtain the following new findings in the context of concurrent data processing in SSDs. (1) Write performance is largely independent of access patterns (regardless of being sequential or random), and can even outperform reads, which is opposite to the long-existing common understanding about slow writes on SSDs. (2) One performance concern comes from interference between concurrent reads and writes, which causes substantial performance degradation. (3) Parallel I/O performance is sensitive to physical data-layout mapping, which is largely not observed without parallelism. (4) Existing application designs optimized for magnetic disks can be suboptimal for running on SSDs with parallelism. Our study is further supported by a group of case studies in database systems as typical data-intensive applications. With these critical findings, we give a set of recommendations to application designers and system architects for exploiting internal parallelism and maximizing the performance potential of SSDs.
Feng Chen 0005, Rubao Lee, Xiaodong Zhang 0001
HPCA1
2011 Hystor: making the best use of solid state drives in high performance storage systems
abstract
With the fast technical improvement, flash memory based Solid State Drives (SSDs) are becoming an important part of the computer storage hierarchy to significantly improve performance and energy efficiency. However, due to its relatively high price and low capacity, a major system research issue to address is on how to make SSDs play their most effective roles in a high-performance storage system in cost- and performance-effective ways.
Feng Chen 0005, David A. Koufaty, Xiaodong Zhang 0001
ICS1
2011 Differentiated storage services
abstract
We propose an I/O classification architecture to close the widening semantic gap between computer systems and storage systems. By classifying I/O, a computer system can request that different classes of data be handled with different storage system policies. Specifically, when a storage system is first initialized, we assign performance policies to predefined classes, such as the filesystem journal. Then, online, we include a classifier with each I/O command (e.g., SCSI), thereby allowing the storage system to enforce the associated policy for each I/O that it receives.
Michael P. Mesnier, Feng Chen 0005, Jason B. Akers
SOSP2
2010 PS-BC: power-saving considerations in design of buffer caches serving heterogeneous storage devices
abstract
Under a replacement policy, existing operating systems identify and maintain most frequently used storage data in buffer caches located in main memory, aiming at low-latency I/O data accesses. However, replacement policies can also strongly affect energy consumptions of various connected storage devices, which has not been a consideration in the design and implementation of buffer cache management. In this paper, we present a system framework for an energy-aware buffer cache replacement, called PS-BC (power-saving buffer cache). By considering several critical factors affecting system energy consumption, PS-BC can effectively improve system energy efficiency, while it is able to flexibly incorporate conventional performance-oriented buffer cache replacement policies for different performance objectives. Our experimental studies based on a trace-driven simulation show that the PS-BC framework embedded with the CLOCK replacement policy can achieve an energy saving rate of up to 32.5% with a minimal overhead for various workloads.
Feng Chen 0005, Xiaodong Zhang 0001
ISLPED1
2009 MCC-DB: Minimizing Cache Conflicts in Multi-core Processors for Databases
abstract
In a typical commercial multi-core processor, the last level cache (LLC) is shared by two or more cores. Existing studies have shown that the shared LLC is beneficial to concurrent query processes with commonly shared data sets. However, the shared LLC can also be a performance bottleneck to concurrent queries, each of which has private data structures, such as a hash table for the widely used hash join operator, causing serious cache conflicts. We show that cache conflicts on multi-core processors can significantly degrade overall database performance. In this paper, we propose a hybrid system method called MCC-DB for accelerating executions of warehouse-style queries, which relies on the DBMS knowledge of data access patterns to minimize LLC conflicts in multi-core systems through an enhanced OS facility of cache partitioning. MCC-DB consists of three components: (1) a cacheaware query optimizer carefully selects query plans in order to balance the numbers of cache-sensitive and cache-insensitive plans; (2) a query execution scheduler makes decisions to co-run queries with an objective of minimizing LLC conflicts; and (3) an enhanced OS kernel facility partitions the shared LLC according to each query's cache capacity need and locality strength. We have implemented MCC-DB by patching the three components in PostgreSQL and Linux kernel. Our intensive measurements on an Intel multi-core system with warehouse-style queries show that MCC-DB can reduce query execution times by up to 33%.
Rubao Lee, Xiaoning Ding, Feng Chen 0005, Qingda Lu, Xiaodong Zhang 0001
Proc. VLDB Endow.3
2008 Caching for bursts (C-Burst): let hard disks sleep well and work energetically
abstract
High energy consumption has become a critical challenge in all kinds of computer systems. Hardware-supported Dynamic Power Management (DPM) provides a mechanism to save disk energy by transitioning an idle disk to a low-power mode. However, the achievable disk energy saving is mainly dependent on the pattern of I/O requests received at the disk. In particular, for a given number of requests, a bursty disk access pattern serves as a foundation for energy optimization. Aggressive prefetching has been used to increase disk access burstiness and extend disk idle intervals, while caching, a critical component in buffer cache management, has not been paid a specific attention. In the absence of cooperation from caching, the attempt to create bursty disk accesses would often be disturbed due to improper replacement decision made by energy unaware caching policies. In this paper, we present the design of a set of comprehensive energy-aware caching schemes, called C-Burst, and its implementation in Linux kernel 2.6.21. Our caching schemes leverage the 'filtering' effect of buffer cache to effectively reshape the disk access stream to a bursty pattern for energy saving. The experiments under various scenarios show that C-Burst schemes can achieve up to 35% disk energy saving with minimal performance loss.
Feng Chen 0005, Xiaodong Zhang 0001
ISLPED1
2007 FlexFetch: A History-Aware Scheme for I/O Energy Saving in Mobile Computing
abstract
Extension of battery lifetime has always been a major issue for mobile computing. While more and more data are involved in mobile computing, energy consumption caused by I/O operations becomes increasingly large. In a pervasive computing environment, the requested data can be stored both on the local disk of a mobile computer by using the hoarding technique, and on the remote server, where data are accessible via wireless communication. Based on the current operational states of local disk (active or standby), the amount of data to be requested (small or large), and currently available wireless bandwidth (strong or weak reception), data access source can be adaptively selected to achieve maximum energy reduction. To this end, we propose a profile-based I/O management scheme, FlexFetch, that is aware of access history and adaptive to current access environment. Our simulation experiments driven by real-life traces demonstrate that the scheme can significantly reduce energy consumption in a mobile computer compared with existing representative schemes.
Feng Chen 0005, Song Jiang 0001, Weisong Shi, Weikuan Yu
ICPP1
2007 DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch
Xiaoning Ding, Song Jiang 0001, Feng Chen 0005, Kei Davis, Xiaodong Zhang 0001
USENIX ATC3
2007 A buffer cache management scheme exploiting both temporal and spatial localities
abstract
On-disk sequentiality of requested blocks, or their spatial locality, is critical to real disk performance where the throughput of access to sequentially-placed disk blocks can be an order of magnitude higher than that of access to randomly-placed blocks. Unfortunately, spatial locality of cached blocks is largely ignored, and only temporal locality is considered in current system buffer cache managements. Thus, disk performance for workloads without dominant sequential accesses can be seriously degraded. To address this problem, we propose a scheme called DULO ( DU al LO cality) which exploits both temporal and spatial localities in the buffer cache management. Leveraging the filtering effect of the buffer cache, DULO can influence the I/O request stream by making the requests passed to the disk more sequential, thus significantly increasing the effectiveness of I/O scheduling and prefetching for disk performance improvements. We have implemented a prototype of DULO in Linux 2.6.11. The implementation shows that DULO can significantly increases disk I/O throughput for real-world applications such as a Web server, TPC benchmark, file system benchmark, and scientific programs. It reduces their execution times by as much as 53%.
Xiaoning Ding, Song Jiang 0001, Feng Chen 0005
ACM Trans. Storage3
2006 SmartSaver: turning flash drive into a disk energy saver for mobile computers
abstract
In a mobile computer the hard disk consumes a considerable amount of energy. Existing dynamic power management policies usually take conservative approaches to save disk energy, and disk energy consumption remains a serious issue. Meanwhile, the flash drive is becoming a must-have portable storage device for almost every laptop user on travel. In this paper, we propose to make another highly desired use of the flash drive --- saving disk energy. This is achieved by using the flash drive as a standby buffer for caching and prefetching disk data. Our design significantly extends disk idle times with careful and deliberate consideration of the particular characteristics of the flash drive. Trace-driven simulations show that up to 41% of disk energy can be saved with a relatively small amount of data written to the flash drive.
Feng Chen 0005, Song Jiang 0001, Xiaodong Zhang 0001
ISLPED1
2005 DULO: An Effective Buffer Cache Management Scheme to Exploit Both Temporal and Spatial Localities
Song Jiang 0001, Xiaoning Ding, Feng Chen 0005, Enhua Tan, Xiaodong Zhang 0001
FAST3
2005 CLOCK-Pro: An Effective Improvement of the CLOCK Replacement
Song Jiang 0001, Feng Chen 0005, Xiaodong Zhang 0001
USENIX ATC, General Track2