VLDB 2026 Research / reviewers in the wild / expert
Bo Mao 0003
dblp:26/3849-3
· DBLP profile ↗
67ranked-venue papers
18as first author
23since 2021 · last 2026
0000-0002-4819-4583ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 61 · 16 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HRAC: A High-Ratio Lossless Compressor for High-Resolution Astronomical DataabstractThis paper proposes HRAC, a novel compressor targeting high-entropy and highresolution astronomical data of both integer and floating-point types. HRAC leverages the distinct characteristics of high-frequency astronomical data across different dimensions: it reads the data along the dimension with the lowest variability and partitions the stream into blocks. For each block, HRAC calculates the mean value of the data after excluding the maximum and minimum, then applies differential coding using this mean against all data within the block to generate a residual sequence. This residual sequence is encoded using a prefix code that combines ideas from ExpGolomb and Elias gamma coding. When compressing floating-point data, the local smoothness assumption crucial for differential prediction is violated if two close values straddle zero. HRAC addresses this by selectively moving the sign bit after the exponent bits. For decompression, several parameters are required per block. Unlike the conventional approach of storing parameters in each block header, HRAC computes the optimal parameters from the previous block and reuses them for the next, eliminating their storage overhead. Experiments on multiple datasets demonstrate that HRAC delivers superior overall performance compared with other compressors, as shown in Fig. 1. Heshan Wang, Jingwen Guo, Suzhen Wu, Bo Mao 0003 |
DCC | 5 |
| 2026 | Xerxes: Extensive Exploration of Scalable Hardware Systems with CXL-Based Simulation Framework
Yuda An, Shushu Yi, Bo Mao 0003, Qiao Li 0001, Mingzhe Zhang 0005, Diyu Zhou, Ke Zhou 0001, Nong Xiao 0001, Guangyu Sun 0003, Yingwei Luo, Jie Zhang 0048 |
FAST | 3 |
| 2025 | SPDK+: Low Latency or High Power Efficiency? We Take BothabstractSPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK. Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048 |
HotStorage | 8 |
| 2025 | Gemina: A Coordinated and High-Performance Memory Deduplication EngineabstractMemory deduplication is widely used to effectively reduce the memory footprint in operating systems, while huge pages are employed to enhance access performance by increasing TLB hit rates. Unfortunately, duplicate huge pages are very rare in memory, leading to most huge pages being split into base pages during deduplication, which degrades access performance by up to $\mathbf{4 3 \%}$. Although asynchronous huge page promotion attempts to elevate multiple base pages to huge pages to maximize performance, deduplication disrupts the uniformity of page attributes, limiting the effectiveness of such promotions. To address this problem, we designed Gemina, a new memory deduplication scheme that coordinates huge page management. Gemina employs a fine-grained detector to distinguish pages, redundancy and hotness, a distributor to process pages conservatively and avoid frequent switches, and an adapter to organize memory for fast conversion. This approach balances the trade-off between performance and space efficiency in memory management. Experimental evaluations show that, compared to the naive KSM-based deduplication scheme, Gemina can achieve $96.1 \%$ of the memory savings of KSM while also improving memory access performance by $\mathbf{4 3. 1 \%}$ in random access. Zhehua Zhang, Suzhen Wu, Wenyan You, Chunfeng Du, Bo Mao 0003 |
HPCA | 5 |
| 2025 | BL-Tree: The Best of Both Worlds by Combining B+- Tree on Top and LSM - Tree on BottomabstractThe shattered and overlapped Level-0 data organization is the primary cause of write stall and read amplification problems in LSM-Tree-based Key-Value (KV) stores: (1) Level-0 to Level-1 compaction involves a large amount of data which induces write stalls, and (2) A point lookup needs to access multiple files in Level-0 which leads to significant read amplification. To address the problem, we propose BL-Tree by replacing the shattered Level-0 in LSM-Tree with a B+-Tree in byte-addressable Persistent Memory (PM). The sorted B+-Tree of Level-0 can accelerate the point lookup speed and reduce read/write amplification. BL-Tree further conducts the locality-aware and parallel compaction from the B+-Tree in PM (Level-0) to the lower levels of LSM-Tree in SSDs by only moving cold data downward, thus alleviating the write stalls and reducing the read/write amplification simultaneously. The extensive experiments on the prototype of BL- Tree show that it definitely avoids the write stalls and significantly reduces the read/write amplification. As a result, BL-Tree reduces the P99 tail latency by 65.2 × than LevelDB-PM and speeds up the throughput by more than 2 × under workloads with spatial locality than other KV stores. Suzhen Wu, Zuocheng Wang, Shengzhe Wang 0001, Jiahong Chen, Chunfeng Du, Ke Zhou 0001, Jie Zhang 0048, Bo Mao 0003 |
ICDE | 8 |
| 2025 | Achieving Better Benefits via Flexible Feature Matching in Post-Deduplication Delta CompressionabstractCloud or distributed storage systems characterized by high data redundancy necessitate effective data reduction techniques to reduce storage costs. Post-deduplication delta compression has proven effective by eliminating both duplicated and similar yet non-duplicated chunks. However, existing approaches often rely on fixed-feature matching for resemblance detection, which, while fast, may lead to lower reduction ratios and not robust benefits across various datasets. In this paper, we introduce BePro, a novel system that integrates Flexible Feature Matching (§IV-A) to achieve better benefits in post-deduplication delta compression. BePro employs Gain Filtering (§IV-B) to identify high-gain chunks while discarding low-gain similar chunks, ensuring robust benefits across different datasets. Additionally, BePro implements a new indexing structure, LSH-Delta (§IV-C), to search for similar chunks and utilizes Index Load Balancer (§IV-D) for efficient resemblance detection by exploiting the distribution characteristics of similar chunks. Furthermore, the Index Manager (§IV-E) skillfully manages memory space overhead, ensuring memory efficiency. We implemented a pipeline prototyping framework to facilitate the evaluation of BePro and other leading techniques. Extensive experiments demonstrate that BePro improves the data-reduction ratios by up to$1.15 \times-2.35 \times$while achieving comparable speed. Fengkui Yang, Bo Mao 0003, Liang Bao, Dongying Zhang, Chunhua Li 0002, Ke Zhou 0001 |
IPDPS | 2 |
| 2025 | HaParallel: Hit Ratio-Aware Parallel Aggressive Eviction Cache Management Algorithm for SSDsabstractSolid-state drives (SSDs) can be classified into two types: with and without built-in cache. In terms of performance, SSDs featuring cache exhibit a substantial performance advantage over their non-cache counterparts. The main focus of this paper is the management of built-in cache in SSDs. From a large number of previous studies, we observe that cache hit ratios remain relatively modest for the majority of workloads. First, based on this observation, we adopt an aggressive eviction policy, diverging from traditional eviction algorithms that follow an on-demand eviction policy. Next, considering the temporal locality and parallelism of cached data, we introduce multi-level linked lists to organize cached data. In this way, a smaller computational load can be used to increase the probability of triggering advanced commands. Finally, drawing inspiration from congestion control algorithm in computer networking, we design a cache hit ratio-aware unit. This unit can employ varying degrees of aggressive eviction policies based on its own state. The aim is to maximize the execution of advanced commands while limiting the impact of the aggressive eviction policy on the hit ratio. Our experimental simulations on real workloads show a substantial improvement in average response time compared to LRU, VBBMS, and Req-block, with our method achieving reductions of 19.7%, 19.4%, and 29.9%, respectively. Liangkuan Su, Mingwei Lin, Bo Mao 0003, Zeshui Xu |
ACM Trans. Storage | 3 |
| 2025 | Analyzing Request Volatility of I/O Temporal Behaviors in Mobile Storage WorkloadsabstractThe design and performance optimization of flash-based storage subsystems are crucial for improving the system performance of Android-based smartphones. However, it highly relies on wisdom derived from mobile storage workload studies of smartphone applications. From the temporal perspective, our burstiness diagnosis reveals that the arrival processes of I/O requests in 33 smartphone applications, are significantly bursty, especially for read requests. This article studies the correlation of inter-arrival times of read and write requests, and compares the correlations for read and write requests in the four same types of mobile applications. We first observe that read requests of 85% of applications and write requests of 88% of applications present a certain degree of correlation over a longer time range. Then, we further conduct Hurst parameter estimation for these mobile applications with mature statistical tools. All estimated Hurst parameters are larger than 0.5, confirming the existence of self-similarity in a majority of mobile application workloads. Finally, we deploy a flexible I/O request generator for smartphone applications based on the parameters measured from actual traces. Experimental results show that the proposed generator can accurately generate request sequences for various mobile applications and more faithfully characterize the heavy-tail properties of I/O activities in mobile storage workloads than traditional models. Qiang Zou 0005, Bo Mao 0003, Suzhen Wu, Yujuan Tan, Donghong Qin |
ACM Trans. Storage | 2 |
| 2024 | LearnedFTL: A Learning-Based Page-Level FTL for Reducing Double Reads in Flash-Based SSDsabstractWe present LearnedFTL, a new on-demand pagelevel flash translation layer (FTL) design, which employs learned indexes to improve the address translation efficiency of flashbased SSDs. The first of its kind, it reduces the number of double reads induced by address translation in random read accesses. LearnedFTL proposes three key techniques: an in-place-update linear model to build learned indexes efficiently, a virtual PPN representation to obtain contiguous PPNs for sorted LPNs, and a group-based allocation and model training via GC/rewrite strategy to reduce the training overhead. By tightly integrating the aforementioned key techniques, LearnedFTL considerably speeds up address translation while reducing the number of flash read accesses caused by the address translation. Our extensive experiments on a FEMU-based prototype show that LearnedFTL can reduce up to 55.5% address translation-induced double reads. As a result, LearnedFTL reduces the P99 tail latency by 2.9× ∼ 12.2× with an average of 5.5× and 8.2× compared to the state-of-the-art TPFTL and LeaFTL schemes, respectively. Shengzhe Wang 0001, Zihang Lin, Suzhen Wu, Hong Jiang 0001, Jie Zhang 0048, Bo Mao 0003 |
HPCA | 6 |
| 2024 | A Practical and Accurate Battery Emulator for Android SmartphonesabstractWe present a novel battery emulation layer for virtual machines, designed to overcome the limitations of traditional battery management and virtualization technologies. This layer enables the testing and simulation of battery-related features, such as low-power operation modes and hibernation, by emulating battery behavior within virtual environments. Unlike existing solutions that depend on host system batteries, our approach allows precise control and simulation of battery characteristics independently from the host system, addressing challenges such as the necessity for physical battery discharges or the lack of batteries in host systems. Additionally, we employs a modular design to virtualize key charging components, including adapters, transformers, and batteries, into distinct models: Adapter, Buck/Boost, and Battery. It supports a wide range of complex battery simulations and enhances testing efficiency by providing accurate and dynamic power information. Our experiments on a Cuttlefish-based prototype show that the system reaches the fitting accuracy of 92.1% on average and up to 98%. It significantly reduces costs, enhances testing efficiency, and ensures safety in the development and research of various battery technologies. Yikun Wu, Ningjian Zhang, Senbin Xu, Baisong Dai, Maoxin Ye, Suzhen Wu, Bo Mao 0003, Qihang Hu, Caiqiang He, Zhipeng Zhong |
HPCC | 7 |
| 2024 | Flagger: Cooperative Acceleration for Large-Scale Cross-Silo Federated Learning AggregationabstractCross-silo federated learning (FL) leverages homomorphic encryption (HE) to obscure the model updates from the clients. However, HE poses the challenges of complex cryptographic computations and inflated ciphertext sizes. As cross-silo FL scales to accommodate larger models and more clients, the overheads of HE can overwhelm a CPU-centric aggregator architecture, including excessive network traffic, enormous data volume, intricate computations, and redundant data movements. Tackling these issues, we propose Flagger, an efficient and high-performance FL aggregator. Flagger meticulously integrates the data processing unit (DPU) with computational storage drives (CSD), employing these two distinct near-data processing (NDP) accelerators as a holistic architecture to collaboratively enhance FL aggregation. With the delicate delegation of complex FL aggregation tasks, we build Flagger-DPU and Flagger-CSD to exploit both in-network and in-storage HE acceleration to streamline FL aggregation. We also implement Flagger-Runtime, a dedicated software layer, to coordinate NDP accelerators and enable direct peer-to-peer data exchanges, markedly reducing data migration burdens. Our evaluation results reveal that Flagger expedites the aggregation in FL training iterations by ${436\%}$ on average, compared with traditional CPU-centric aggregators. Xiurui Pan, Yuda An, Shengwen Liang, Bo Mao 0003, Mingzhe Zhang 0005, Qiao Li 0001, Myoungsoo Jung, Jie Zhang 0048 |
ISCA | 4 |
| 2024 | SSD Failures in Large-Scale Data Centers: What? Why? and How?abstractWith SSDs gradually replacing HDDs as the main-stream storage media in modern large-scale data centers, SSD failure analysis has become increasingly important. We conducted an in-depth data-driven analysis of the failure characteristics of an SSD-based data center in Alibaba based on the failure datasets in 2018 and 2019 and SMART logs on December 31, 2019. Our objectives focus on 3 “W”s, illustrated as follows. What factors influence the occurrence of failures? Why do these factors influence the occurrence of failures? How can companies reduce the occurrence of SSD failures? We expect that our findings and analysis can benefit future SSD-based storage system designs. Wenyan You, Jiayuan Dong, Xingdi Feng, Zeyun Chen, Bo Mao 0003, Suzhen Wu |
NAS | 5 |
| 2024 | ScalaAFA: Constructing User-Space All-Flash Array Engine with Holistic Designs
Shushu Yi, Xiurui Pan, Qiao Li 0001, Chenxi Wang 0005, Bo Mao 0003, Myoungsoo Jung, Jie Zhang 0048 |
USENIX ATC | 6 |
| 2024 | FSDedup: Feature-Aware and Selective Deduplication for Improving Performance of Encrypted Non-Volatile Main MemoryabstractEnhancing the endurance, performance, and energy efficiency of encrypted Non-Volatile Main Memory (NVMM) can be achieved by minimizing written data through inline deduplication. However, existing approaches applying inline deduplication to encrypted NVMM suffer from substantial performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). In the preliminary ESD [ 14 ], we proposed the Error Correcting Code (ECC) assisted selective deduplication scheme, utilizing the ECC information as a fingerprint to identify similar data effectively and then leveraging the selective deduplication technique to eliminate a large amount of redundant data with high reference counts. In this article, we proposed FSDedup. Compared with ESD, FSDedup could leverage the prefetch cache to reduce the read overhead during similarity comparison and utilize the cache refresh mechanism to identify further and eliminate more redundant data. Extensive experimental evaluations demonstrate that FSDedup can enhance the performance of the NVMM system further than the ESD. Experimental results show that FSDedup can improve both write and read speed by up to 1.8×, enhance Instructions Per Cycle by up to 1.5×, and reduce energy consumption by up to 2.0×, compared to ESD. Chunfeng Du, Zihang Lin, Suzhen Wu, Yifei Chen 0011, Jiapeng Wu, Shengzhe Wang 0001, Weichun Wang 0002, Bo Mao 0003 |
ACM Trans. Storage | 9 |
| 2023 | iKnowFirst: An Efficient DPU-Assisted Compaction for LSM-Tree-Based Key-Value StoresabstractIn scenarios with write-intensive workloads, LSM-tree-based key-value stores, such as RocksDB, suffer from compaction-induced performance degradation. RocksDB provides configurable compaction options to mitigate the severe read/write amplification problems associated with compaction. The advent of the Data Processing Unit (DPU) allows us to better utilize the configurable options of RocksDB to guide the key-value store system in choosing a suitable compaction strategy with prior knowledge of the workload characteristics. This paper proposes iKnowFirst, an efficient DPU-assisted key-value store. iKnowFirst (1) sets a data buffer on the DPU and separates hot-cold data to relieve the pressure of subsequent LSM-tree compaction, (2) senses the characteristics of the workloads in advance, and dynamically guides RocksDB to choose different compaction modes or enable/disable compaction when the workloads change, to cope with the scenario of write outbreak, and (3) implements an auto-selecting interface for compaction strategies selection. Our prototype implementation and experimental results show that iKnowFirst achieves 3.2× improvement compared to the original RocksDB on write-intensive and highly skewed workloads while showing acceptable performance under read-intensive workloads. Jiahong Chen, Shengzhe Wang 0001, Suzhen Wu, Bo Mao 0003 |
ASAP | 5 |
| 2023 | ESD: An ECC-assisted and Selective Deduplication for Encrypted Non-Volatile Main MemoryabstractReducing write data to encrypted Non-Volatile Main Memory (NVMM) can directly improve NVMM’s endurance, performance, and energy efficiency. However, existing works that straightforwardly apply inline deduplication on encrypted NVMM can significantly lead to system performance degradation due to high computing, memory footprint, and index-lookup overhead to generate, store, and query the cryptographic hash (fingerprint). This paper proposes ESD, an ECC-assisted and Selective Deduplication for encrypted NVMM by exploiting both the device characteristics (ECC mechanism) and the workload characteristics (content locality). First, ESD utilizes the ECC information associated with each cache line evicted from the Last-Level Cache (LLC) as the fingerprint to identify data similarity and avoids the costly hash calculating overhead on the non-duplicate cache lines. Second, ESD leverages selective deduplication to exploit the content locality within cache lines by only storing the fingerprints with high reference counts in the memory cache to reduce the memory space overhead and avoid fingerprints NVMM_lookup operations. The experimental results show that ESD can significantly speed up the writes by up to 3.4x, 4.3x, and 2.6x, speed up the reads by up to 5.3x, 5.0x, and 2.0x, and reduce the energy consumption by up to 96.3%, 96.2%, and 56.6% than Baseline, Dedup SHA1, and DeWrite, respectively. Meanwhile, ESD also can significantly outperform other schemes in tail latency. Chunfeng Du, Suzhen Wu, Jiapeng Wu, Bo Mao 0003, Shengzhe Wang 0001 |
HPCA | 4 |
| 2023 | LearnedSync: A Learning-Based Sync Optimization for Cloud Storage
Suzhen Wu, Shengzhe Wang 0001, Chunfeng Du, Jiayang Guo, Yijie Pan, Naian Xiao, Bo Mao 0003 |
ICA3PP (2) | 8 |
| 2023 | EaD: ECC-Assisted Deduplication With High Performance and Low Memory Overhead for Ultra-Low Latency Flash StorageabstractData deduplication has become a commodity feature in flash storage products to effectively reduce redundant write data and improve space efficiency. However, it also introduces computing and memory overhead to generate and store the cryptographic hash (fingerprint) in face of the moderate data redundancy in primary storage. With the advent of 3D XPoint and Z-NAND technologies, and the stronger cryptographic hash functions in use, such as SHA-256, both the computing and memory overheads are increasingly serious performance bottlenecks for inline data deduplication in these ultra-low latency flash storage. To address these problems, we propose an ECC-assisted Deduplication approach, called EaD, which exploits the ECC property and the asymmetric read-write performance characteristics of modern flash storage. EaD first identifies data similarity by leveraging the device-generated ECC values of data chunks as their fingerprints, significantly reducing the costly MD5/SHA-based cryptographic hash computing and alleviating the memory space overhead. Based on the identification results, similar data chunks and their ECCs are read from the flash to perform a byte-by-byte comparison in memory to definitively identify and remove redundant data chunks. Our experiments show that the EaD approach significantly increases I/O performance by up to 4.2${\times }$, with an average of 2.5${\times }$, compared with the existing MD5/SHA- and sampling-based deduplication approaches. Suzhen Wu, Chunfeng Du, Weidong Zhu 0002, Jindong Zhou, Hong Jiang 0001, Bo Mao 0003, Lingfang Zeng |
IEEE Trans. Computers | 6 |
| 2023 | FASTSync: A FAST Delta Sync Scheme for Encrypted Cloud Storage in High-bandwidth Network EnvironmentsabstractMore and more data are stored in cloud storage, which brings two major challenges. First, the modified files in the cloud should be quickly synchronized to ensure data consistency, e.g., delta synchronization (sync) achieves efficient cloud sync by synchronizing only the updated part of the file. Second, the huge data in the cloud needs to be deduplicated and encrypted, e.g., Message-Locked Encryption (MLE) implements data deduplication by encrypting the content among different users. However, when combined, a few updates in the content can cause large sync traffic amplification for both keys and ciphertext in the MLE-based cloud storage, significantly degrading the cloud sync efficiency. A feature-based encryption sync scheme, FeatureSync, is proposed to address the delta amplification problem. However, with further improvement of the network bandwidth, the performance of FeatureSync stagnates. In our preliminary experimental evaluations, we find that the bottleneck of the computational overhead in the high-bandwidth network environments is the main bottleneck in FeatureSync. In this article, we propose an enhanced feature-based encryption sync scheme FASTSync to optimize the performance of FeatureSync in high-bandwidth network environments. The performance evaluations on a lightweight prototype implementation of FASTSync show that FASTSync reduces the cloud sync time by 70.3% and the encryption time by 37.3%, on average, compared with FeatureSync. Suzhen Wu, Zhanhong Tu, Zuocheng Wang, Zhirong Shen, Wei Wang 0424, Weichun Wang 0002, Bo Mao 0003 |
ACM Trans. Storage | 9 |
| 2022 | DedupHR: Exploiting Content Locality to Alleviate Read/Write Interference in Deduplication-Based Flash StorageabstractThe read/write interference problem of flash storage, where on-going write requests prevent read requests from accessing available data and significantly increase read latency, remains a critical concern with the rapid evolutions of flash chips from SLC through MLC to TLC and on to QLC, with PLC on the horizon, especially under workloads with a mixture of read and write requests. To significantly improve the read performance in face of read/write interference, which is often more critical than the write performance, we propose an enhanced Hot Read Data Replication scheme for deduplication-based flash storage, called DedupHR, to alleviate the read/write interference problem. DedupHR exploits the content locality in the deduplication-based flash storage and outsources the data blocks with high reference count to a surrogate space such as a dedicated spare flash chip or an over-provisioned space in an SSD. By servicing some conflicted read requests on the surrogate flash space, DedupHR can significantly alleviate, if not entirely eliminate, the contention between the read requests and the on-going write requests. The evaluation results show that DedupHR significantly improves the state-of-the-art schemes in terms of the system performance and cost efficiency. Consequently, the tail-latency of the deduplication-based flash storage is also reduced. Suzhen Wu, Chunfeng Du, Bo Mao 0003, Hong Jiang 0001 |
IEEE Trans. Computers | 4 |
| 2021 | SimiEncode: A Similarity-based Encoding Scheme to Improve Performance and Lifetime of Non-Volatile Main MemoryabstractNon-Volatile Memories (NVMs) have shown tremendous potential to be the next generation of main memory, yet they are still seriously hampered by the high write latency and limited endurance. In this paper, we first unveil via realworld benchmark analysis that the words within the same cache line showcase a high degree of similarity. We therefore present SimiEncode, a low-overhead and effective Similarity-based Encoding approach. SimiEncode relieves writes to NVMs by (1) generating a mask word with minimized differences to the words within a cache line, (2) encoding each word with the associated mask word by simple XOR operations, and (3) writing a single tag bit to indicate the resulting zero word after encoding. Our prototype implementation of SimiEncode and extensive evaluations driven by 15 state-of-the-art benchmarks demonstrate that, compared with existing approaches, SimiEncode significantly prolongs the lifetime and improves the performance. Importantly, SimiEncode is orthogonal to and can be easily incorporated into existing bit flipping optimizations. Suzhen Wu, Jiapeng Wu, Zhirong Shen, Zuocheng Wang, Bo Mao 0003 |
ICCD | 6 |
| 2021 | When Delta Sync Meets Message-Locked Encryption: a Feature-based Delta Sync Scheme for Encrypted Cloud StorageabstractAs increasingly prevalent, more and more data are stored in the cloud storage, which brings us two major challenges. First, the modified files in the cloud should be quickly synchronized (sync) to ensure data consistency, e.g., delta sync achieves efficient cloud sync by synchronizing only the updated part of the file. Second, the huge data in the cloud needs to be deduplicated and encrypted, e.g., message-locked encryption (MLE) implements data deduplication by encrypting the content between different users. However, when both are combined, few updates in the content can cause large sync traffic amplification for both keys and ciphertext in the MLE-based cloud storage, which significantly degrading the cloud sync efficiency. In this paper, we propose an feature-based encryption sync scheme FeatureSync to improve the performance of synchronizing multiple encrypted files by merging several files before synchronizing. The performance evaluations on a lightweight prototype implementation of FeatureSync show that FeatureSync reduces the cloud sync time by 72.6% and the cloud sync traffic by 78.5% on average, compared with the state-of-the-art sync schemes. Suzhen Wu, Zhanhong Tu, Zuocheng Wang, Zhirong Shen, Bo Mao 0003 |
ICDCS | 5 |
| 2021 | CAGC: A Content-aware Garbage Collection Scheme for Ultra-Low Latency Flash-based SSDsabstractWith the advent of new flash-based memory technologies with ultra-low latency, directly applying inline data deduplication in flash-based storage devices can degrade the system performance since key deduplication operations lie on the shortened critical write path of such devices. To address the problem, we propose a Content-Aware Garbage Collection scheme (CAGC), which embeds the data deduplication into the data movement workflow of the Garbage Collection (GC) process in ultra-low latency flash-based SSDs. By parallelizing the operations of valid data pages migration, hash computing and flash block erase, the deduplication-induced performance overhead is alleviated and redundant page writes during the GC period are eliminated. To further reduce data writes and write amplification during GC, CAGC separates and stores data pages in different regions based on their reference counts. The performance evaluation of our CAGC prototype implemented in FlashSim shows that CAGC significantly reduces the number of flash blocks erased and data pages migrated during GC, leading to improved user I/O performance and reliability of ultra-low latency flash-based SSDs. Suzhen Wu, Chunfeng Du, Hong Jiang 0001, Zhirong Shen, Bo Mao 0003 |
IPDPS | 6 |
| 2020 | EaD: a Collision-free and High Performance Deduplication Scheme for Flash Storage SystemsabstractInline deduplication is a popular technique to effectively reduce the write traffic and improve the space efficiency for flash-based storage. However, it also introduces computing and memory overhead to generate and store the cryptographic hash (fingerprint). Along the advent of 3D XPoint and Z-NAND technologies with vastly improved latency and bandwidth, both the computing and memory overheads are becoming much more pronounced in deduplication-based flash storage with cryptographic hash functions in use. To address these problems, we propose an ECC (Error Correcting Code) assisted deduplication approach, called EaD, which exploits the ECC property and the asymmetric read-write performance characteristics of modern flash-based storage. EaD first identifies data similarity based on the fingerprints of data chunks represented by their ECC values, thus significantly reducing the costly cryptographic hash computing and alleviating the memory space overhead. Based on the identification results, similar data chunks and their ECCs are read from the flash to perform a byte-by-byte comparison in memory to definitively identify and remove redundant data chunks. Our experiments show that the EaD approach significantly reduces the I/O latency by an average of 1.92× and 1.86×, and reduces the memory consumption by an average of 35.0% and 21.9%, compared with the existing SHA- and sampling-based deduplication approaches, respectively. Suzhen Wu, Jindong Zhou, Weidong Zhu 0002, Hong Jiang 0001, Zhirong Shen, Bo Mao 0003 |
ICCD | 7 |
| 2020 | GC-Steering: GC-Aware Request Steering and Parallel Reconstruction Optimizations for SSD-Based RAIDsabstractSolid-state disk (SSD)-based redundant array of independent disks (RAIDs) have been widely deployed in high-end enterprize systems to provide high-performance and highly reliable storage for data-intensive computing. However, SSD-based RAIDs suffer from significant performance degradation whenever user I/O requests conflict with the ongoing garbage collection (GC) operations which introduce tail latency. Moreover, the performance characteristics of SSDs make the traditional HDD-based RAID reconstruction algorithms are not compatible with or suitable for SSD-based RAIDs. In this article, we proposed GC-aware request steering (GC-Steering), a scheme aware of the GC process within an SSD-based RAID, to significantly boost the performance and reliability of SSD-based RAIDs. GC-Steering effectively outsources the popular read requests and all write requests addressed to the SSD currently in the GC state to a staging space, such as a dedicated spare SSD or the reserved space of each SSD within the RAID. GC-Steering also accelerates the performance of the failure-recovery process by both request steering and parallel recovery. Our extensive evaluations on a lightweight GC-Steering prototype driven by HPC-like and real-world enterprize workloads show that the GC-Steering scheme significantly reduces the average response time by an average of 63.3% and 65.8%, compared with the state-of-the-art local GC and global GC schemes. Moreover, the GC-Steering scheme also significantly reduces the average response times by an average of 62.3% during RAID reconstruction than the normal state. Suzhen Wu, Weidong Zhu 0002, Yingxin Han, Hong Jiang 0001, Bo Mao 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | HotR: Alleviating Read/Write Interference with Hot Read Data Replication for Flash StorageabstractThe read/write interference problem of flash storage remains a critical concern under workloads with a mixture of read and write requests. To significantly improve the read performance in face of read/write interference, we propose a Hot Data Replication scheme for flash storage, called HotR. HotR utilizes the asymmetric read and write performance characteristics of flash-based SSDs and outsources the popular read data to a surrogate space such as a dedicated spare flash chip or an over-provisioned space within an SSD. By servicing some conflicted read requests on the surrogate flash space, HotR can alleviate, if not entirely eliminate, the contention between the read requests and the on-going write requests. The evaluation results show that HotR improves the state-of-the-art scheme in the system performance and cost efficiency significantly. Consequently, the tail-latency of the flash-based storage systems is also reduced. Suzhen Wu, Bo Mao 0003, Hong Jiang 0001 |
DATE | 3 |
| 2019 | PandaSync: Network and Workload Aware Hybrid Cloud Sync OptimizationabstractWith the widespread use and increasing popularity of cloud storage, more and more data are moved to the cloud, making cloud storage a platform for both data sharing among users, devices, and data backup for data reliability. Thus, it is critically important to ensure data consistency through efficient cloud synchronization (sync). The existing cloud synchronization schemes are either delta sync, which sends only the updated portion of a file but incurs high compute overhead of data deduplication for small files, or full sync, which avoids data deduplication by sending the full file but wastes network bandwidth and lengthens sync time by transferring significant amount of redundant data over the networks for large files. In this paper, we propose a hybrid cloud sync scheme, PandaSync, that combines full sync and delta sync dynamically based on file size and network conditions. To further improve small-file sync performance, we propose an optimization, Full2Sync, that merges the sync request with the file-sending request to reduce the number of network round-trips between the client and the cloud servers. The experiments conducted on our lightweight prototype implementation of PandaSync show that PandaSync reduces the sync time by an average of 85.1% and 74.6% from the delta sync scheme and full sync scheme, respectively. Suzhen Wu, Longquan Liu, Hong Jiang 0001, Hao Che, Bo Mao 0003 |
ICDCS | 5 |
| 2019 | IOFollow: Improving the performance of VM live storage migration with IO following in the cloud
Bo Mao 0003, Yaodong Yang 0003, Suzhen Wu, Hong Jiang 0001, Kuanching Li |
Future Gener. Comput. Syst. | 1 |
| 2019 | Mitigating and Tolerating Read Disturbance in STT-MRAM-Based Main Memory via Device and Architecture InnovationsabstractAs an important nonvolatile memory technology, spin transfer torque magnetoresistive RAM (STT-MRAM) is widely considered as a universal memory solution for future processors. Employing STT-MRAM as the main memory offers a wide variety of benefits, but also results in unique design challenges. In particular, read disturbance characterizes accidental data corruption in STT-MRAM after it is read, leading to the need of restoring data back to memory after each read operation. In this paper, we propose both device and architecture innovations to mitigate and tolerate read disturbance. First, we quantitatively demonstrate the relationship between read disturbance and key device parameters, conducting a number of read disturbance mitigation schemes. These device-level schemes turn out to be effective in reducing the read disturbance probability, but come with costs on other design metrics. Consequently, we further propose a restore-aware memory controller design at the architecture level to tolerate read disturbance. Since the extra restores incurred by read disturbance greatly change the timing scenarios that conventional memory controllers were optimized for, directly adopting restore-agnostic DRAM memory management techniques will lead to suboptimal designs for STT-MRAM. Therefore, we propose restore-aware policy selection (RAPS), a dynamic and hybrid row buffer management scheme that factors in the inevitable data restores in STT-MRAM-based main memory. RAPS monitors the row buffer hit rate at run time, dynamically switching between two static page-closure policies. By factoring in restores, RAPS accurately captures the optimal design points, achieving optimal policy selections at run time. Our experimental results show that RAPS significantly improves system performance and energy efficiency compared to conventional page-closure policies. Armin Haj Aboutalebi, Ethan C. Ahn, Bo Mao 0003, Suzhen Wu, Lide Duan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Improving Flash Memory Performance and Reliability for Smartphones With I/O DeduplicationabstractFlash-based storage subsystem is the key component that affects the system performance, reliability, and cost efficiency of Android-based smartphones. In this paper, we first introduce a trace collection tool specifically designed to capture the I/O requests with important content features in Android-based smartphones, which are critically important but rarely available in content-aware designs and optimizations, such as JProbe and Netlink. Based on the analysis of the traces collected from 15 popular mobile applications, we find that 20%-40% of the I/O requests on the I/O critical path of the storage stack are redundant and this data redundancy is minimally shared among different applications. Based on this key observation, we propose a content-aware optimization, called APP-Dedupe, that applies data deduplication on the I/O critical path to improve both performance and efficiency by reducing write amplification and improving GC efficiency of the flash storage on Android smartphones. The evaluation results show that APP-Dedupe reduces the GC overhead by an average of 41.5%, reduces the response times by up to 15.4% and reduces the amount of write data by an average of 45.2%. Bo Mao 0003, Jindong Zhou, Suzhen Wu, Hong Jiang 0001, Weijian Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Overcome the GC-Induced Performance Variability in SSD-Based RAIDs With Request RedirectionabstractThe I/O bottleneck has become an increasingly daunting challenge for big data analytics along with the explosive growth in data volume. Flash-based SSDs become promising to replace the hard disk drives. However, garbage collection (GC) operations in SSDs have a significant impact on the SSD performance, thus leading to performance variability in SSD-based RAIDs. To address this problem, we propose request redirection (RR) by exploiting the asymmetric read-write performance characteristics of SSDs and the hot-spare SSD in SSD-based RAIDs to alleviate the GC-induced performance variability. RR services the incoming read requests to the SSD currently in GC state by reconstructing the read data from other SSDs in the same stripe within SSD-based RAIDs. For the incoming write data to the SSD in the GC state, RR temporarily stores the write data on the hot-spare SSD and concurrently updates the corresponding parity in the SSD-based RAIDs. Extensive evaluations on the RR prototype show that the RR scheme significantly reduces the average response time and alleviates the performance variability, compared with the local GC and global GC schemes. Suzhen Wu, Bo Mao 0003, Xiaoxi Chen, Kuanching Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | PFP: Improving the Reliability of Deduplication-based Storage Systems with Per-File ParityabstractData deduplication weakens the reliability of storage systems since by design it removes duplicate data chunks common to different files and forces these files to share a single physical date chunk, or critical chunk, after deduplication. Thus, the loss of a single such critical data chunk can potentially render all referencing (sharing) files unavailable. However, the reliability issue in deduplication-based storage systems has not received adequate attention. Existing approaches introduce data redundancy after files have been deduplicated, either by replication on critical data chunks, i.e., chunks with high reference count, or RAID schemes on unique data chunks, which means that these schemes are based on individual unique data chunks rather than individual files. This can leave individual files vulnerable to losses, particularly in the presence of transient and unrecoverable data chunk errors such as latent sector errors. To address this file reliability issue, this paper proposes a Per-File Parity (short for PFP) scheme to improve the reliability of deduplication-based storage systems. PFP computes the XOR parity within parity groups of data chunks of each file after the chunking process but before the data chunks are deduplicated. Therefore, PFP can provide parity redundancy protection for all files by intra-file recovery and a higher-level protection for data chunks with high reference counts by inter-file recovery. Our reliability analysis and extensive data-driven, failure-injection based experiments conducted on a prototype implementation of PFP show that PFP significantly outperforms the existing redundancy solutions, DTR and RCR, in system reliability, tolerating multiple data chunk failures and guaranteeing file availability upon multiple data chunk failures. Moreover, a performance evaluation shows that PFP only incurs an average of 5.7 percent performance degradation to the deduplication-based storage system. Suzhen Wu, Bo Mao 0003, Hong Jiang 0001, Huagao Luan, Jindong Zhou |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | GC-Aware Request Steering with Improved Performance and Reliability for SSD-Based RAIDsabstractSSD-based RAIDs have been widely deployed in high-end enterprise systems to provide high-performance and highly reliable storage for data-intensive computing. However, SSD-based RAIDs suffer from significant performance degradation whenever user I/O requests conflict with the ongoing Garbage Collection (GC) operations which introduces tail latency. Moreover, the performance characteristics of SSDs make the traditional HDD-based RAID reconstruction algorithms are not compatible with or suitable for SSD-based RAIDs. In this paper, we proposed GC-aware Request Steering (short for GC-Steering), a scheme aware of the GC process within an SSD-based RAID, to significantly boost the performance and reliability of SSD-based RAIDs. GC-Steering effectively outsources the popular read requests and all write requests addressed to the SSD currently in the GC state to a staging space such as a dedicated spare SSD or the reserved space of each SSD within the RAID. GC-Steering also accelerates the performance of the failure-recovery process by both request steering and parallel recovery. Our extensive evaluations on a lightweight GC-Steering prototype driven by HPC-like and real-world enterprise workloads show that the GC-Steering scheme significantly reduces the average response time by an average of 63.3% and 65.8%, compared with the state-of-the-art LGC and GGC schemes. Moreover, GC-Steering scheme also significantly reduces the average response times by an average of 55.7% during RAID reconstruction than the normal state. Suzhen Wu, Weidong Zhu 0002, Guixin Liu, Hong Jiang 0001, Bo Mao 0003 |
IPDPS | 5 |
| 2018 | Improving Reliability of Deduplication-Based Storage Systems with Per-File ParityabstractThe reliability issue in deduplication-based storage systems has not received adequate attention. Existing approaches introduce data redundancy after files have been deduplicated, either by replication on critical data chunks, i.e., chunks with high reference count, or RAID schemes on unique data chunks, which means that these schemes are based on individual unique data chunks rather than individual files. This can leave individual files vulnerable to losses, particularly in the presence of transient and unrecoverable data chunk errors such as latent sector errors. To address this file reliability issue, this paper proposes a Per-File Parity (short for PFP) scheme to improve the reliability of deduplication-based storage systems. PFP computes the XOR parity within parity groups of data chunks of each file after the chunking process but before the data chunks are deduplicated. Therefore, PFP can provide parity redundancy protection for all files by intra-file recovery and a higher-level protection for data chunks with high reference counts by inter-file recovery. Our reliability analysis and extensive data-driven, failure-injection based experiments conducted on a prototype implementation of PFP show that PFP significantly outperforms the existing redundancy solutions, DTR and RCR, in system reliability, tolerating multiple data chunk failures and guaranteeing file availability upon multiple data chunk failures. Moreover, a performance evaluation shows that PFP only incurs an average of 5.7% performance degradation to the deduplication-based storage system. Suzhen Wu, Huagao Luan, Bo Mao 0003, Hong Jiang 0001, Gen Niu, Hui Rao, Jindong Zhou |
SRDS | 3 |
| 2018 | PP: Popularity-based Proactive Data Recovery for HDFS RAID systems
Suzhen Wu, Weidong Zhu 0002, Bo Mao 0003, Kuanching Li |
Future Gener. Comput. Syst. | 3 |
| 2018 | Improving the SSD Performance by Exploiting Request Characteristics and Internal ParallelismabstractWith the explosive growth in the data volume, the I/O bottleneck has become an increasingly daunting challenge for big data analytics. It is urgent and important to introduce high-performance flash-based solid state drives (SSDs) into the storage systems. However, since the existing systems are primarily designed for conventional magnetic hard disk drives, directly incorporating SSDs in the existing systems cannot fully exploit SSDs' performance advantages. In this paper, we propose a new I/O scheduler for SSDs, namely Amphibian, that exploits the high-level request characteristics and low-level parallelism of flash chips to improve the performance of SSD-based storage systems. Amphibian includes two performance enhancement schemes: 1) size-based request ordering, which prioritizes requests with small sizes in processing and 2) garbage collection (GC)-aware request dispatching that delays issuing requests to flash chips that are in the GC state. These two schemes employed in Amphibian significantly reduce the average waiting times of the requests from the host. Our extensive evaluation results derived from three types of SSDs show that, compared with the existing I/O schedulers, Amphibian greatly improves both throughput and average response times for SSD-based storage systems, thus improving the I/O performance of the systems. Bo Mao 0003, Suzhen Wu, Lide Duan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | EDC: Improving the Performance and Space Efficiency of Flash-Based Storage Systems with Elastic Data CompressionabstractBy leveraging data reduction technologies, such as data compression, all flash-based storage systems can have the same total cost of ownership (TCO) as traditional HDD-based storage systems. Thus, data compression has become a commodity feature for space efficiency and reliability in flash-based storage systems by reducing write traffic and space capacity demand. However, it introduces noticeable processing overheads on the critical I/O path, which degrades the system performance significantly. Existing data compression schemes for flash-based storage systems use fixed compression algorithms for all the incoming write data, failing to recognize and exploit the significant diversity in compressibility and access patterns of data and missing an opportunity to improve the system performance, the space efficiency or both. To achieve a reasonable trade-off between these two important design objectives, in this paper we introduce an Elastic Data Compression scheme, called EDC, which exploits the data compressibility and access intensity characteristics by judiciously matching data of different compressibility with different compression algorithms while leveraging the access idleness. Specifically, for compressible data blocks EDC exploits the compression diversity of the workload, and employs algorithms of higher compression rate in periods of lower system utilization and algorithms of lower compression rate in periods of higher system utilization. For non-compressible (or very lowly compressible) data blocks, it will write them through to the flash storage directly without any compression. The experiments conducted on our lightweight prototype implementation of the EDC system show that EDC saves storage space by up to 38.7 percent, with an average of 33.7 percent . In addition, it significantly outperforms the fixed compression schemes in the I/O performance measure by up to 61.4 percent, with an average of 36.7 percent. Bo Mao 0003, Suzhen Wu, Hong Jiang 0001, Yaodong Yang 0003, Zaifa Xi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | SnapMig: Accelerating VM Live Storage Migration by Leveraging the Existing VM Snapshots in the CloudabstractVirtual Machine (VM) live storage migration is becoming increasingly important and indispensable in the current cloud data centers, for the purposes of load balance, hardware maintenance and system upgrade. Nevertheless, conventional VM migration approaches induce significant extra storage and network traffic to the source server that is already heavily loaded or scheduled for upgrade or repair. As a result, both the VM performance perceived by the application/user and the migration performance are degraded significantly. In this paper, we aim to address this problem by proposing a novel scheme, called SnapMig, to improve the VM live storage migration efficiency and eliminate its performance impact on user applications at the source server by effectively leveraging the existing VM snapshots in backup servers. By outsourcing the task of transferring VM base image and snapshots to the destination server to backup servers, the source server only needs to migrate the latest state changes to the destination server, leading to simultaneous improvement on VM performance, migration time and multiple-VM migration efficiency. Our lightweight prototype implementation of the SnapMig scheme demonstrates that, compared with the state-of-the-art approaches, SnagMig can significantly reduce the migration time and improve the source-server VM performance at the same time. Moreover, the performance improvement provided by SnapMig becomes much more pronounced with multiple concurrent VM migrations. Yaodong Yang 0003, Bo Mao 0003, Hong Jiang 0001, Yuekun Yang, Hao Luo 0009, Suzhen Wu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Parallelism and Garbage Collection Aware I/O Scheduler with Improved SSD PerformanceabstractIn this paper, we propose PGIS, a parallelism and garbage collection aware I/O Scheduler, which identifies the hot data based on trace characteristics to exploit the channel level internal parallelism of flash-based storage systems. PGIS not only fully exploits abundant channel resource in the SSD, but also it introduces a hot data identification mechanism to reduce the garbage collection overhead. By dispatching hot read data to different channel, the channel level internal parallelism is fully exploited. By dispatching hot write data to the same physical block, the garbage collection overhead has been alleviated. The experiment results show that compared with existing I/O schedulers, PGIS improves the response time and garbage collection performance significantly. Consequently, PGIS reduces the garbage collection overhead up to 30.9%, while exploiting channel level internal parallelism. Jiayang Guo, Yiming Hu, Bo Mao 0003, Suzhen Wu |
IPDPS | 3 |
| 2017 | Elastic Data Compression with Improved Performance and Space Efficiency for Flash-Based Storage SystemsabstractData compression has become a commodity feature for space efficiency and reliability in flash-based storage systems by reducing write traffic and space capacity demand. However, it introduces noticeable processing overheads on the critical I/O path, which degrades the system performance significantly. Existing data compression schemes for flash-based storage systems use fixed compression algorithms for all the incoming write data, failing to recognize and exploit the significant diversity in compressibility and access patterns of data and missing an opportunity to improve the system performance, the space efficiency or both. To achieve a reasonable trade-off between these two important design objectives, in this paper we introduce an Elastic Data Compression scheme, called EDC, which exploits the data compressibility and access intensity characteristics by judiciously matching data of different compressibility with different compression algorithms while leveraging the access idleness. Specifically, for compressible data blocks EDC exploits the compression diversity of the workload, and employs algorithms of higher compression rate in periods of lower system utilization and algorithms of lower compression rate in periods of higher system utilization. For non-compressible (or very lowly compressible) data blocks, it will write them through to the flash storage directly without any compression. The experiments conducted on our lightweight prototype implementation of the EDC system show that EDC saves storage space by up to 38.7%, with an average of 33.7%. In addition, it significantly outperforms the fixed compression schemes in the I/O performance measure by up to 61.4%, with an average of 36.7%. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Yaodong Yang 0003, Zaifa Xi |
IPDPS | 1 |
| 2017 | DAC: Improving storage availability with Deduplication-Assisted Cloud-of-Clouds
Suzhen Wu, Kuanching Li, Bo Mao 0003, Minghong Liao |
Future Gener. Comput. Syst. | 3 |
| 2017 | Improving Performance for Flash-Based Storage Systems through GC-Aware Cache ManagementabstractFlash-based SSDs have been extensively deployed in modern storage systems to satisfy the increasing demand of storage performance and energy efficiency. However, Garbage Collection (GC) is an important performance concern for flash-based SSDs, because it tends to disrupt the normal operations of an SSD. This problem continues to plague flash-based storage systems, particularly in the high performance computing and enterprise environment. An important root cause for this problem, as revealed by previous studies, is the serious contention for the flash resources and the severe mutually adversary interference between the user I/O requests and GC-induced I/O requests. The on-board buffer cache within SSDs serves to play an essential role in smoothing the gap between the upper-level applications and the lower-level flash chips and alleviating this problem to some extend. Nevertheless, the existing cache replacement algorithms are well optimized to reduce the miss rate of the buffer cache by reducing the I/O traffic to the flash chips as much as possible, but without considering the GC operations within the flash chips. Consequently, they fail to address the root cause of the problem and thus are far from being sufficient and effective in reducing the expensive I/O traffic to the flash chips that are in the GC state. To address this important performance issue in flash-based storage systems, particularly in the HPC and enterprise environment, we propose a Garbage Collection aware Replacement policy, called GCaR, to improve the performance of flash-based SSDs. The basic idea is to give higher priority to caching the data blocks belonging to the flash chips that are in the GC state. This substantially lessens the contentions between the user I/O operations and the GC-induced I/O operations. To verify the effectiveness of GCaR, we have integrated it into the SSD extended Disksim simulator. The simulation results show that GCaR can significantly improve the storage performance by up to 40.7 percent in terms of the average response times. Suzhen Wu, Bo Mao 0003, Yanping Lin, Hong Jiang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Exploiting the Data Redundancy Locality to Improve the Performance of Deduplication-Based Storage SystemsabstractThe chunk-lookup disk bottleneck and the read amplification problems are two great challenges for deduplication-based storage systems and restrict the applicability of data deduplication for large-scale data volumes. Previous studies and our experimental evaluations have shown that the amount of redundant data shared among different types of applications is negligible. Based on the observations, we propose AA-Plus which effectively groups the hash index of the same application together and divides the whole hash index into different groups based on the application types. Moreover, it groups the data chunks of the same application together on the disks. The extensive trace-driven experiments conducted on our lightweight prototype implementation of AA-Plus show that compared with AA-Dedupe, AA-Plus significantly speeds up the write throughput by a factor of up to 6.9 and with an average of 3.1, and speeds up the read throughput by a factor of up to 3.3 and with an average of 1.9. Suzhen Wu, Bo Mao 0003 |
ICPADS | 3 |
| 2016 | GCaR: Garbage Collection aware Cache Management with Improved Performance for Flash-based SSDsabstractGarbage Collection (GC) is an important performance concern for flash-based SSDs, because it tends to disrupt the normal operations of an SSD. This problem continues to plague flash-based storage systems, particularly in the high performance computing and enterprise environment. An important root cause for this problem, as revealed by previous studies, is the serious contention for the flash resources and the severe mutually adversary interference between the user I/O requests and GC-induced I/O requests. The on-board buffer cache within SSDs serves to play an essential role in smoothing the gap between the upper-level applications and the lower-level flash chips and alleviating this problem to some extend. Nevertheless, the existing cache replacement algorithms are well optimized to reduce the miss rate of the buffer cache by reducing the I/O traffic to the flash chips as much as possible, but without considering the GC operations within the flash chips. Consequently, they fail to address the root cause of the problem and thus are far from being sufficient and effective in reducing the expensive I/O traffic to the flash chips that are in the GC state. Suzhen Wu, Yanping Lin, Bo Mao 0003, Hong Jiang 0001 |
ICS | 3 |
| 2016 | Leveraging Data Deduplication to Improve the Performance of Primary Storage Systems in the CloudabstractWith the explosive growth in data volume, the I/O bottleneck has become an increasingly daunting challenge for big data analytics in the Cloud. Recent studies have shown that moderate to high data redundancy clearly exists in primary storage systems in the Cloud. Our experimental studies reveal that data redundancy exhibits a much higher level of intensity on the I/O path than that on disks due to relatively high temporal access locality associated with small I/O requests to redundant data. Moreover, directly applying data deduplication to primary storage systems in the Cloud will likely cause space contention in memory and data fragmentation on disks. Based on these observations, we propose a performance-oriented I/O deduplication, called POD, rather than a capacity-oriented I/O deduplication, exemplified by iDedup, to improve the I/O performance of primary storage systems in the Cloud without sacrificing capacity savings of the latter. POD takes a two-pronged approach to improving the performance of primary storage systems and minimizing performance overhead of deduplication, namely, a request-based selective deduplication technique, called Select-Dedupe, to alleviate the data fragmentation and an adaptive memory management scheme, called iCache, to ease the memory contention between the bursty read traffic and the bursty write traffic. We have implemented a prototype of POD as a module in the Linux operating system. The experiments conducted on our lightweight prototype implementation of POD show that POD significantly outperforms iDedup in the I/O performance measure by up to 87.9 percent with an average of 58.8 percent. Moreover, our evaluation results also show that POD achieves comparable or better capacity savings than iDedup. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Lei Tian 0001 |
IEEE Trans. Computers | 1 |
| 2016 | LDM: Log Disk Mirroring with Improved Performance and Reliability for SSD-Based Disk ArraysabstractWith the explosive growth in data volume, the I/O bottleneck has become an increasingly daunting challenge for big data analytics. Economic forces, driven by the desire to introduce flash-based Solid-State Drives (SSDs) into the high-end storage market, have resulted in hybrid storage systems in the cloud. However, a single flash-based SSD cannot satisfy the performance, reliability, and capacity requirements of enterprise or HPC storage systems in the cloud. While an array of SSDs organized in a RAID structure, such as RAID5, provides the potential for high storage capacity and bandwidth, reliability and performance problems will likely result from the parity update operations. In this article, we propose a Log Disk Mirroring scheme (LDM) to improve the performance and reliability of SSD-based disk arrays. LDM is a hybrid disk array architecture that consists of several SSDs and two hard disk drives (HDDs). In an LDM array, the two HDDs are mirrored as a write buffer that temporally absorbs the small write requests. The small and random write data are written on the mirroring buffer by using the logging technique that sequentially appends new data. The small write data are merged and destaged to the SSD-based disk array during the system idle periods. Our prototype implementation of the LDM array and the performance evaluations show that the LDM array significantly outperforms the pure SSD-based disk arrays by a factor of 20.4 on average, and outperforms HPDA by a factor of 5.0 on average. The reliability analysis shows that the MTTDL of the LDM array is 2.7 times and 1.7 times better than that of pure SSD-based disk arrays and HPDA disk arrays. Suzhen Wu, Bo Mao 0003, Xiaolan Chen, Hong Jiang 0001 |
ACM Trans. Storage | 2 |
| 2016 | Exploiting Workload Characteristics and Service Diversity to Improve the Availability of Cloud Storage SystemsabstractWith the increasing utilization and popularity of the cloud infrastructure, more and more data are moved to the cloud storage systems. This makes the availability of cloud storage services critically important, particularly given the fact that outages of cloud storage services have indeed happened from time to time. Thus, solely depending on a single cloud storage provider for storage services can risk violating the service-level agreement (SLA) due to the weakening of service availability. This has led to the notion of Cloud-of-Clouds, where data redundancy is introduced to distribute data among multiple independent cloud storage providers, to address the problem. The key in the effectiveness of the Cloud-of-Clouds approaches lies in how the data redundancy is incorporated and distributed among the clouds. However, the existing Cloud-of-Clouds approaches utilize either replication or erasure codes to redundantly distribute data across multiple clouds, thus incurring either high space or high performance overheads. In this paper, we propose a hybrid redundant data distribution approach, called HyRD, to improve the cloud storage availability in Cloud-of-Clouds by exploiting the workload characteristics and the diversity of cloud providers. In HyRD, large files are distributed in multiple cost-efficient cloud storage providers with erasure-coded data redundancy while small files and file system metadata are replicated on multiple high-performance cloud storage providers. The experiments conducted on our lightweight prototype implementation of HyRD show that HyRD improves the cost efficiency by 33.4 and 20.4 percent, and reduces the access latency by 58.7 and 34.8 percent than the DuraCloud and RACS schemes, respectively. Bo Mao 0003, Suzhen Wu, Hong Jiang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | WAIO: Improving Virtual Machine Live Storage Migration for the Cloud by Workload-Aware IO OutsourcingabstractVirtual Machine (VM) live storage migration is widely performed in the current cloud data centers, for the purposes of load balance, hardware maintenance and system upgrade. Nevertheless, conventional migration approaches, such as Dirty Block Tracking (DBT), do not address the problem of IO interference between VM IO requests and Migration IO requests during the migration period, which degrades both the VM IO performance and the migration performance. In this paper, we propose a Workload-Aware IO Outsourcing scheme, short for WAIO, to improve the VM live storage migration efficiency. WAIO effectively outsources the VM's working set to a surrogate device during the migration and creates separate IO path for servicing the VM IO requests. By outsourcing VM IO requests from the original storage to the surrogate device, the VM live storage migration process can be performed on the original storage, no longer interfered, while the outsourced VM IO requests are serviced separately and thus much more quickly. Our lightweight prototype implementation of WAIO and extensive trace-driven experiments demonstrate that, compared with the existing migration approach DBT, WAIO significantly improves the VM's IO performance during the migration process. Moreover, WAIO allows the hypervisor to migrate a VM at a higher migration speed, without sacrificing the VM's IO performance. Yaodong Yang 0003, Hong Jiang 0001, Bo Mao 0003, Lei Tian 0001, Yuekun Yang, Junjie Qian |
CloudCom | 3 |
| 2015 | MC-RAIS: Multi-chunk Redundant Array of Independent SSDs with Improved Performance
Suzhen Wu, Weijian Yang, Bo Mao 0003, Yanping Lin |
ICA3PP (4) | 3 |
| 2015 | Exploiting request characteristics and internal parallelism to improve SSD performanceabstractIn this paper, we propose a new I/O scheduler for SSDs, called Amphibian, which exploits the up-level request characteristics and the low-level internal parallelism of flash chips to improve the performance of SSD-based storage systems. Amphibian includes two parts: the size-based request ordering that gives higher priority to first processing the small requests and the Garbage Collection (GC) aware request dispatching that avoids issuing requests to the flash chips that are in the GC state. By first processing the small requests and avoiding issuing the GC-conflict requests in the I/O waiting queue, the average waiting times of the requests are reduced significantly. The extensive evaluation results show that compared with existing I/O schedulers, Amphibian improves the throughput and the average response times significantly. Consequently, the I/O performance of the SSD-based storage systems is improved. Bo Mao 0003, Suzhen Wu |
ICCD | 1 |
| 2015 | Improving Storage Availability in Cloud-of-Clouds with Hybrid Redundant Data DistributionabstractWith the increasing utilization and popularity of the cloud infrastructure, more and more data are moved to the cloud storage systems. This makes the availability of cloud storage services critically important, particularly given the fact that outages of cloud storage services have indeed happened from time to time. Thus, solely depending on a single cloud storage provider for storage services can risk violating the service-level agreement (SLA) due to the weakening of service availability. This has led to the notion of Cloud-of-Clouds, where data redundancy is introduced to distribute data among multiple independent cloud storage providers, to address the problem. The key in the effectiveness of the Cloud-of-Clouds approaches lies in how the data redundancy is incorporated and distributed among the clouds. However, the existing Cloud-of-Clouds approaches utilize either replication or erasure codes to redundantly distribute data across multiple clouds, thus incurring either high space or high performance overheads. In this paper, we propose a hybrid redundant data distribution approach, called HyRD, to improve the cloud storage availability in Cloud-of-Clouds by exploiting the workload characteristics and the diversity of cloud providers. In HyRD, large files are distributed in multiple cost-efficient cloud storage providers with erasure-coded data redundancy while small files and file system metadata are replicated on multiple high-performance cloud storage providers. The experiments conducted on our lightweight prototype implementation of HyRD show that HyRD improves the cost efficiency by 33.4% and 20.4%, and reduces the access latency by 58.7% and 34.8% than the DuraCloud and RACS schemes, respectively. Bo Mao 0003, Suzhen Wu, Hong Jiang 0001 |
IPDPS | 1 |
| 2015 | Proactive Data Migration for Improved Storage Availability in Large-Scale Data CentersabstractIn face of high partial and complete disk failure rates and untimely system crashes, the executions of low-priority background tasks become increasingly frequent in large-scale data centers. However, the existing algorithms are all reactive optimizations and only exploit the temporal locality of workloads to reduce the user I/O requests during the low-priority background tasks. To address the problem, this paper proposes Intelligent Data Outsourcing (IDO), a zone-based and proactive data migration optimization, to significantly improve the efficiency of the low-priority background tasks. The main idea of IDO is to proactively identify the hot data zones of RAID-structured storage systems in the normal operational state. By leveraging the prediction tools to identify the upcoming events, IDO proactively migrates the data blocks belonging to the hot data zones on the degraded device to a surrogate RAID set in the large-scale data centers. Upon a disk failure or crash reboot, most user I/O requests addressed to the degraded RAID set can be serviced directly by the surrogate RAID set rather than the much slower degraded RAID set. Consequently, the performance of the background tasks and user I/O performance during the background tasks are improved simultaneously. Our lightweight prototype implementation of IDO and extensive trace-driven experiments on two case studies demonstrate that, compared with the existing state-of-the-art approaches, IDO effectively improves the performance of the low-priority background tasks. Moreover, IDO is portable and can be easily incorporated into any existing algorithms for RAID-structured storage systems. Suzhen Wu, Hong Jiang 0001, Bo Mao 0003 |
IEEE Trans. Computers | 3 |
| 2014 | Exploiting Content Locality to Improve the Performance and Reliability of Phase Change Memory
Suzhen Wu, Zaifa Xi, Bo Mao 0003, Hong Jiang 0001 |
ICA3PP (2) | 3 |
| 2014 | POD: Performance Oriented I/O Deduplication for Primary Storage Systems in the CloudabstractRecent studies have shown that moderate to high data redundancy clearly exists in primary storage systems in the Cloud. Our experimental studies reveal that data redundancy exhibits a much higher level of intensity on the I/O path than that on disks due to the relatively high temporal access locality associated with small I/O requests to redundant data. On the other hand, we also observe that directly applying data deduplication to primary storage systems in the Cloud will likely cause space contention in memory and data fragmentation on disks. Based on these observations, we propose a Performance-Oriented I/O Deduplication approach, called POD, rather than a capacity-oriented I/O deduplication approach, represented by iDedup, to improve the I/O performance of primary storage systems in the Cloud without sacrificing capacity savings of the latter. The salient feature of POD is its focus on not only the capacity-sensitive large writes and files, as in iDedup, but also the performance-sensitive while capacity-insensitive small writes and files. The experiments conducted on our lightweight prototype implementation of POD show that POD significantly outperforms iDedup in the I/O performance measure by up to 87.9% with an average of 58.8%. Moreover, our evaluation results also show that POD achieves comparable or better capacity savings than iDedup. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Lei Tian 0001 |
IPDPS | 1 |
| 2014 | Read-Performance Optimization for Deduplication-Based Storage Systems in the CloudabstractData deduplication has been demonstrated to be an effective technique in reducing the total data transferred over the network and the storage space in cloud backup, archiving, and primary storage systems, such as VM (virtual machine) platforms. However, the performance of restore operations from a deduplicated backup can be significantly lower than that without deduplication. The main reason lies in the fact that a file or block is split into multiple small data chunks that are often located in different disks after deduplication, which can cause a subsequent read operation to invoke many disk IOs involving multiple disks and thus degrade the read performance significantly. While this problem has been by and large ignored in the literature thus far, we argue that the time is ripe for us to pay significant attention to it in light of the emerging cloud storage applications and the increasing popularity of the VM platform in the cloud. This is because, in a cloud storage or VM environment, a simple read request on the client side may translate into a restore operation if the data to be read or a VM suspended by the user was previously deduplicated when written to the cloud or the VM storage server, a likely scenario considering the network bandwidth and storage capacity concerns in such an environment. To address this problem, in this article, we propose SAR, an SSD (solid-state drive)-Assisted Read scheme, that effectively exploits the high random-read performance properties of SSDs and the unique data-sharing characteristic of deduplication-based storage systems by storing in SSDs the unique data chunks with high reference count, small size, and nonsequential characteristics. In this way, many read requests to HDDs are replaced by read requests to SSDs, thus significantly improving the read performance of the deduplication-based storage systems in the cloud. The extensive trace-driven and VM restore evaluations on the prototype implementation of SAR show that SAR outperforms the traditional deduplication-based and flash-based cache schemes significantly, in terms of the average response times. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Yinjin Fu, Lei Tian 0001 |
ACM Trans. Storage | 1 |
| 2013 | Leveraging data deduplication to improve the performance of primary storage systems in the cloudabstractRecent studies have shown that moderate to high data redundancy exists in primary storage systems, such as VM-based, enterprise and HPC storage systems, which indicates that the data deduplication technology can be used to effectively reduce the write traffic and storage space in such environments. However, our experimental studies reveal that applying data deduplication to primary storage systems will cause space contention in main memory and data fragmentation on disks. This is in part because applying data deduplication introduces significant index memory overhead to the existing system and in part because a file or block is split into multiple small data chunks that are often located in non-sequential locations on disks after deduplication. This fragmentation of data can cause a subsequent read operation to invoke many disk I/O requests, thus leading to performance degradation. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Lei Tian 0001 |
SoCC | 1 |
| 2012 | HerpRap: A Hybrid Array Architecture Providing Any Point-in-Time Data Tracking for DatacenterabstractBoth physical disk failure and logical errors such as software error, user abuse and virus attacks may cause data lose. The risk of logical errors is far greater than physical disk failure. Moreover, existing RAID solution cannot satisfy the reliability requirement in face of the logical errors in data centers. It is therefore becoming increasingly important for RAID-based storage systems to be able to recover data to any point-in-time when logical errors occur. We proposed a novel storage array architecture, Herp Rap, which is able to recover data from both physical disk failure and logical errors. We have implemented a prototype of Herp Rap and carried out extensive performance measurements using DBT-2 and file system benchmarks. Our experiments demonstrated that the proposed Herp Rap is able to track or recover data to any point-in-time quickly by tracing back the history of block logs. Moreover, Herp Rap outperforms existing HDD-based or SSD-based RAID5 with copy-on-write (COW) snapshot in terms of performance, energy efficiency, failure recovery ability and reliability. Lingfang Zeng, Dan Feng 0001, Bo Mao 0003, Jianxi Chen, Qingsong Wei, Wenguo Liu 0004 |
CLUSTER | 3 |
| 2012 | IDO: Intelligent Data Outsourcing with Improved RAID Reconstruction Performance in Large-Scale Data Centers
Suzhen Wu, Hong Jiang 0001, Bo Mao 0003 |
LISA | 3 |
| 2012 | SAR: SSD Assisted Restore Optimization for Deduplication-Based Storage Systems in the CloudabstractThe explosive growth of digital content results in enormous strains on the storage systems in the cloud environment. The data deduplication technology has been demonstrated to be very effective in shortening the backup window and saving the network bandwidth and storage space in cloud backup, archiving and primary storage systems such as VM platforms. However, the delay and power consumption of the restore operations from a deduplicated storage can be significantly higher than those without deduplication. The main reason lies in the fact that a file or block is split into multiple small data chunks that are often located in non-sequential locations on HDDs after deduplication, which can cause a subsequent read operation to invoke many HDD I/O requests involving multiple disk seeks. To address this problem, in this paper we propose SAR, an SSD Assisted Restore scheme, that effectively exploits the high random-read performance and low power-consumption properties of SSDs and the unique data sharing characteristic of deduplication-based storage system by storing in SSDs the unique data chunks with high reference count, small size and non-sequential characteristics. In this way, many critical random-read requests to HDDs are replaced by read requests to SSDs, thus significantly improving the system performance and energy efficiency. The extensive trace-driven and VM restore evaluations on the prototype implementation of SAR show that SAR outperforms the traditional deduplication-based schemes significantly, in terms of both restore performance and energy efficiency. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Yinjin Fu, Lei Tian 0001 |
NAS | 1 |
| 2012 | HPDA: A hybrid parity-based disk array for enhanced performance and reliabilityabstractFlash-based Solid State Drive (SSD) has been productively shipped and deployed in large scale storage systems. However, a single flash-based SSD cannot satisfy the capacity, performance and reliability requirements of the modern storage systems that support increasingly demanding data-intensive computing applications. Applying RAID schemes to SSDs to meet these requirements, while a logical and viable solution, faces many challenges. In this article, we propose a Hybrid Parity-based Disk Array architecture (short for HPDA), which combines a group of SSDs and two hard disk drives (HDDs) to improve the performance and reliability of SSD-based storage systems. In HPDA, the SSDs (data disks) and part of one HDD (parity disk) compose a RAID4 disk array. Meanwhile, a second HDD and the free space of the parity disk are mirrored to form a RAID1-style write buffer that temporarily absorbs the small write requests and acts as a surrogate set during recovery when a disk fails. The write data is reclaimed to the data disks during the lightly loaded or idle periods of the system. Reliability analysis shows that the reliability of HPDA, in terms of MTTDL (Mean Time To Data Loss), is better than that of either pure HDD-based or SSD-based disk array. Our prototype implementation of HPDA and the performance evaluations show that HPDA significantly outperforms either HDD-based or SSD-based disk array. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Lei Tian 0001, Dan Feng 0001, Jianxi Chen, Lingfang Zeng |
ACM Trans. Storage | 1 |
| 2011 | Beyond mirroring: multi-version disk arraywith improved performance and energy efficiencyabstractPerformance and power consumption are two important design objectives for data centers consisting of thousands or tens of thousands of disks (or disk arrays). To leverage the two objectives, in this study we propose a multi-version disk array (MDA). The main idea of MDA is to exploit the I/O workload characteristics to guide the replication strategy by replicating multiple versions of the popular data blocks and simply offloading the write data to the free space of the reserved version region, thus achieving high performance in the burst period and low power consumption in the idle period. Our prototype implementation of MDA and the performance evaluations show that the performance of MDA outperforms that of traditional RAID10 by up to 34.4% and 42.3% in terms of the average response time for the online transaction processing (OLTP) application I/O and search engine I/O, respectively. Moreover, the energy efficiency of MDA outperforms that of RAID10 by up to 48.7% and 36.4%, respective to the aforementioned measures. Bo Mao 0003, Suzhen Wu, Dan Feng 0001 |
J. Zhejiang Univ. Sci. C | 1 |
| 2011 | Improving Availability of RAID-Structured Storage Systems by Workload OutsourcingabstractDue to the contention for the shared disk bandwidth, the user I/O intensity can significantly impact the performance of the online low-priority background tasks, thus reducing the reliability and availability of RAID-structured storage systems. In this paper, we propose a novel and practical scheme, called WorkOut (I/O Workload Outsourcing), to significantly boost the performance of those low-priority background tasks. WorkOut effectively outsources all write requests and popular read requests originally targeted at the degraded RAID set that is performing the low-priority background tasks to a surrogate RAID set. The lightweight prototype implementation of WorkOut and extensive trace-driven and benchmark-driven experiments on two case studies demonstrate that, compared with existing approaches, WorkOut effectively improves the performance of the low-priority background tasks, such as RAID reconstruction and RAID resynchronization. Importantly, WorkOut is portable and can be easily incorporated into any existing optimizing algorithms for RAID-structured storage systems. Suzhen Wu, Hong Jiang 0001, Dan Feng 0001, Lei Tian 0001, Bo Mao 0003 |
IEEE Trans. Computers | 5 |
| 2010 | HPDA: A hybrid parity-based disk array for enhanced performance and reliabilityabstractA single flash-based Solid State Drive (SSD) can not satisfy the capacity, performance and reliability requirements of a modern storage system supporting increasingly demanding data-intensive computing applications. Applying RAID schemes to SSDs to meet these requirements, while a logical and viable solution, faces many challenges. In this paper, we propose a Hybrid Parity-based Disk Array architecture, HPDA, which combines a group of SSDs and two hard disk drives (HDDs) to improve the performance and reliability of SSD-based storage systems. In HPDA, the SSDs (data disks) and part of one HDD (parity disk) compose a RAID4 disk array. Meanwhile, a second HDD and the free space of the parity disk are mirrored to form a RAID1-style write buffer that temporarily absorbs the small write requests and acts as a surrogate set during recovery when a disk fails. The write data is reclaimed back to the data disks during the lightly loaded or idle periods of the system. Reliability analysis shows that the reliability of HPDA, in terms of MTTDL (Mean Time To Data Loss), is better than that of either pure HDD-based or SSD-based disk array. Our prototype implementation of HPDA and performance evaluations show that HPDA significantly outperforms either HDD-based or SSD-based disk array. Bo Mao 0003, Hong Jiang 0001, Dan Feng 0001, Suzhen Wu, Jianxi Chen, Lingfang Zeng, Lei Tian 0001 |
IPDPS | 1 |
| 2009 | WorkOut: I/O Workload Outsourcing for Boosting RAID Reconstruction Performance
Suzhen Wu, Hong Jiang 0001, Dan Feng 0001, Lei Tian 0001, Bo Mao 0003 |
FAST | 5 |
| 2009 | JOR: A Journal-guided Reconstruction Optimization for RAID-Structured Storage SystemsabstractThis paper proposes a simple and practical RAID reconstruction optimization scheme, called JOurnal-guided Reconstruction (JOR). JOR exploits the fact that significant portions of data blocks in typical disk arrays are unused. JOR monitors the storage space utilization status at the block level to guide the reconstruction process so that only failed data on the used stripes is recovered to the spare disk. In JOR, data consistency is ensured by the requirement that all blocks in a disk array be initialized to zero (written with value zero) during synchronization while all blocks in the spare disk also be initialized to zero in the background. JOR can be easily incorporated into any existing reconstruction approach to optimize it, because the former is independent of and orthogonal to the latter. Experimental results obtained from our JOR prototype implementation demonstrate that JOR reduces reconstruction times of two state-of-the-art reconstruction schemes by an amount that is approximately proportional to the percentage of unused storage space while ensuring data consistency. Suzhen Wu, Dan Feng 0001, Hong Jiang 0001, Bo Mao 0003, Lingfang Zeng, Jianxi Chen |
ICPADS | 4 |
| 2008 | Performance-Directed iSCSI Security with Parallel EncryptionabstractISCSI has been paid much attention since it allows the storage data to be transported over the popular TCP/IP networks and takes advantage of the networks. Carrying data over the TCP/IP networks introduces the security problem. The IP security mechanisms, such as IPSec, always degrade the performance greatly. In this paper we introduced a parallel encryption method, combined with the IPSec A,H to provide the security and high performance for the iSCSI storage system. We have implemented parallel encryption with IPSec AH in the UNH iSCSI software. Numerical results using popular benchmark program and OLTP traces have shown dramatic performance gain and the average response time reduced greatly. Bo Mao 0003, Dan Feng 0001, Suzhen Wu, Jianxi Chen, Lingfang Zeng |
AINA | 1 |
| 2008 | GRAID: A Green RAID Storage Architecture with Improved Energy Efficiency and Reliability
Bo Mao 0003, Dan Feng 0001, Suzhen Wu, Lingfang Zeng, Jianxi Chen, Hong Jiang 0001 |
MASCOTS | 1 |