EDBT 2026 Demo / reviewers in the wild / expert
Jiguang Wan 0001
dblp:91/8380-1
· DBLP profile ↗
86ranked-venue papers
9as first author
45since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 8 first-author · 30 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesabstractStreaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries can be issued at arbitrary timepoints. Existing solutions relying on fixed-size memory or naive compression often suffer from context loss or memory overflow, limiting their effectiveness in long-form, real-time scenarios.We present Vista, a novel framework for scene-aware streaming video QA that enables efficient and scalable reasoning over continuous video streams. The innovation of Vista can be summarized in three aspects: (1) Scene-aware segmentation. Vista dynamically clusters incoming frames into temporally and visually coherent scene units. (2) Scene-aware compression. Each scene is compressed into a compact token representation and stored in GPU memory for efficient index-based retrieval, while the full-resolution frames are offloaded to CPU memory. (3) Scene-aware recall. Upon receiving a question, relevant scenes are selectively recalled and reintegrated into the model’s input space, enabling both efficiency and completeness. Vista is model-agnostic and integrates seamlessly with a variety of vision-language backbones, enabling long-context reasoning without compromising latency or memory efficiency. Extensive experiments on StreamingBench demonstrate that Vista achieves state-of-the-art performance, establishing a strong baseline for real-world streaming video understanding. Haocheng Lu, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
AAAI | 6 |
| 2026 | CEMU: Enabling Full-System Emulation of Computational Storage Beyond Hardware LimitsabstractComputational storage drives (CSDs) present a promising approach to improve system performance through near data processing in SSDs. However, current research platforms are fragmented and inadequate to explore the full design space of CSD systems. Existing hardware and emulator platforms are constrained by physical compute resources, while simulators lack full-system fidelity. To address the problems, we introduce CEMU, a new software-based CSD emulation platform that enables full-system research. It consists of a CSD device emulator and a CSD-oriented software stack. Through a novel virtual machine freezing mechanism, CSD emulation achieves high configurability. While the CSD can utilize the host CPU to physically perform computation to preserve full-system behaviors, the computational delay can be modeled separately to emulate CSDs with CPU-unbounded high computing power. The software stack is designed with two principles, adhering to recent industry CSD standards and being compatible with the existing I/O stack, which is achieved via a newly developed file system FDMFS. We verify CEMU's emulation fidelity across a range of applications by benchmarking against actual CSD hardware, demonstrating average end-to-end performance accuracy of 95% or higher. We also use two case studies on large language model training and LevelDB to demonstrate that CEMU is effective in exploring CSD system research and can uncover insights that have not been discovered in previous research platforms. Jiapin Wang, You Zhou 0009, Kai Lu 0002, Jiguang Wan 0001, Fei Wu 0005, Tao Lu 0014 |
ASPLOS (2) | 6 |
| 2026 | HeapKV: Enabling Efficient Garbage Collection for KV-Separated LSM Stores on Modern SSDsabstractKey-value (KV) separation has emerged as a pivotal solution to tackle write amplification in LSM-tree-based KV stores (LSM stores). However, the garbage collection (GC) mechanism essential for reclaiming obsolete values introduces substantial overheads: (i) Additional LSM-tree I/O operations degrade foreground performance and cause data inconsistency issues; (ii) Exacerbated write amplification arises from excessive valid data migration in update-intensive workloads. Moreover, existing GC optimization schemes fundamentally struggle to balance space overhead, write amplification, and system performance. In this article, we propose HeapKV, a high-performance KV-separated LSM store that improves GC efficiency through three key technologies: (i) A lightweight two-level index and a global garbage view decouple GC operations of value storage from the LSM-tree, eliminating additional I/O operations; (ii) A novel valid data migration scheme mitigates write amplification during space reclamation by in-place overwrites and logical data copying; (iii) SSD-conscious I/O optimizations featuring asynchronous value flushing, fast read paths and concurrent prefetching for range queries. Extensive experiments demonstrate that HeapKV achieves 40%–7.4× higher throughput under diverse workloads with lower write/space amplification, compared to other state-of-the-art KV-separated LSM stores. Kai Lu 0002, Yuanhui Zhou, Nengjie Wang, Jiguang Wan 0001, Bisheng Huang |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | RUNA: Object-Level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal RepresentationsabstractEnabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do not receive supervisory signals from unfamiliar data, leading to overly confident predictions regarding OOD objects. Despite previous progress that estimates OOD uncertainty based on the detection model and in-distribution (ID) samples, we explore using pre-trained vision-language representations for object-level OOD detection. We first discuss the limitations of applying image-level CLIP-based OOD detection methods to object-level scenarios. Building upon these insights, we propose RUNA, a novel framework that leverages a dual encoder architecture to capture rich contextual information and employs a regional uncertainty alignment mechanism to distinguish ID from OOD objects effectively. We introduce a few-shot fine-tuning approach that aligns region-level semantic representations to further improve the model's capability to discriminate between similar objects. Our experiments show that RUNA substantially surpasses state-of-the-art methods in object-level OOD detection, particularly in challenging scenarios with diverse and complex object instances. Jinggang Chen, Xiaoyang Qu, Guokuan Li, Kai Lu 0002, Jiguang Wan 0001, Jing Xiao 0006, Jianzong Wang |
AAAI | 6 |
| 2025 | MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware ExpertsabstractOne of the primary challenges in optimizing large language models (LLMs) for long-context inference lies in the high memory consumption of the Key-Value (KV) cache. Existing approaches, such as quantization, have demonstrated promising results in reducing memory usage. However, current quantization methods cannot take both effectiveness and efficiency into account. In this paper, we propose MoQAE, a novel mixed-precision quantization method via mixture of quantization-aware experts. First, we view different quantization bit-width configurations as experts and use the traditional mixture of experts (MoE) method to select the optimal configuration. To avoid the inefficiency caused by inputting tokens one by one into the router in the traditional MoE method, we input the tokens into the router chunk by chunk. Second, we design a lightweight router-only fine-tuning process to train MoQAE with a comprehensive loss to learn the trade-off between model accuracy and memory usage. Finally, we introduce a routing freezing (RF) and a routing sharing (RS) mechanism to further reduce the inference overhead. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms state-of-the-art KV cache quantization approaches in both efficiency and effectiveness. Haocheng Lu, Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Jianzong Wang |
ACL (1) | 6 |
| 2025 | Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM InferenceabstractRecently, large language models (LLMs) have been able to handle longer and longer contexts. However, a context that is too long may cause intolerant inference latency and GPU memory usage. Existing methods propose mixed-precision quantization to the key-value (KV) cache in LLMs based on token granularity, which is time-consuming in the search process and hardware inefficient during computation. This paper introduces a novel approach called Cocktail, which employs chunk-adaptive mixed-precision quantization to optimize the KV cache. Cocktail consists of two modules: chunk-level quantization search and chunk-level KV cache computation. Chunk-level quantization search determines the optimal bitwidth configuration of the KV cache chunks quickly based on the similarity scores between the corresponding context chunks and the query, maintaining the model accuracy. Furthermore, chunk-level KV cache computation reorders the KV cache chunks before quantization, avoiding the hardware inefficiency caused by mixed-precision quantization in inference computation. Extensive experiments demonstrate that Cocktail outperforms state-of-the-art KV cache quantization methods on various models and datasets. Our code is presented on https://github.com/Sullivan12138/Cocktail. Xiaoyang Qu, Jiguang Wan 0001, Jianzong Wang |
DATE | 4 |
| 2025 | PointActionCLIP: Preventing Transfer Degradation in Point Cloud Action Recognition with a Triple-Path CLIPabstractDirectly applying CLIP to point cloud action recognition can cause severe accuracy collapse. In this paper, we propose PointActionCLIP, which successfully prevents this transfer degradation with a triplepath CLIP, including the image path, the sequence path, and the label path. Specifically, the image path projects the 3D point cloud sequence onto a 2D image sequence and uses a visual encoder to extract its feature. It also captures the temporal feature of the image sequence with a temporal encoding transformer. The sequence path adopts a pretrained sequence encoder to encode the original point cloud sequence to obtain its spatiotemporal feature. The label path encodes the candidate labels with a text encoder. Finally, we fuse the output of the three paths to obtain the predicted action label. Extensive experiments validate that PointActionCLIP outperforms state-of-the-art (SOTA) methods. Shenglin He, Xiaoyang Qu, Jiguang Wan 0001, Jianzong Wang |
ICASSP | 4 |
| 2025 | VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD DetectionabstractAs object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot object-level out-of-distribution (OOD) detection arises. This task becomes crucial in ensuring the reliability of detectors in open-world settings. While existing methods have demonstrated success in image-level OOD detection using pre-trained vision-language models like CLIP, directly applying such models to object-level OOD detection presents challenges due to the loss of contextual information and reliance on image-level alignment. To tackle these challenges, we introduce a new method that leverages visual prompts and text-augmented in-distribution (ID) space construction to adapt CLIP for zero-shot object-level OOD detection. Our method preserves critical contextual information and improves the ability to differentiate between ID and OOD objects, achieving competitive performance across different benchmarks. Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
ICASSP | 4 |
| 2025 | MADLLM: Multivariate Anomaly Detection via Pre-trained LLMsabstractWhen applying pre-trained large language models (LLMs) to address anomaly detection tasks, the multivariate time series (MTS) modality of anomaly detection does not align with the text modality of LLMs. Existing methods simply transform the MTS data into multiple univariate time series sequences, which can cause many problems. This paper introduces MADLLM, a novel multivariate anomaly detection method via pre-trained LLMs. We design a new triple encoding technique to align the MTS modality with the text modality of LLMs. Specifically, this technique integrates the traditional patch embedding method with two novel embedding approaches: (i) Skip Embedding, which alters the order of patch processing in traditional methods to help LLMs retain knowledge of previous features, and (ii) Feature Embedding, which leverages contrastive learning to allow the model to better understand the correlations between different features. Experimental results demonstrate that our method outperforms state-of-the-art methods in various public anomaly detection datasets. Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Guokuan Li, Jianzong Wang |
ICME | 4 |
| 2025 | BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic SegmentationabstractSince the point cloud data is inherently irregular and unstructured, point cloud semantic segmentation has always been a challenging task. The graph-based method attempts to model the irregular point cloud by representing it as a graph; however, this approach incurs substantial computational cost due to the necessity of constructing a graph for every point within a large-scale point cloud. In this paper, we observe that boundary points possess more intricate spatial structural information and develop a novel graph attention network known as the Boundary-Aware Graph attention Network (BAGNet). On one hand, BAGNet contains a boundary-aware graph attention layer (BAGLayer), which employs edge vertex fusion and attention coefficients to capture features of boundary points, reducing the computation time. On the other hand, BAGNet employs a lightweight attention pooling layer to extract the global feature of the point cloud to maintain model accuracy. Extensive experiments on standard datasets demonstrate that BAGNet outperforms state-of-the-art methods in point cloud semantic segmentation with higher accuracy and less inference time. Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Shenglin He, Jianzong Wang |
IJCNN | 4 |
| 2025 | DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data Processing
Chen Ding 0012, Sicen Li, Kai Lu 0002, Ting Yao 0001, Daohui Wang, Huatao Wu, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001 |
USENIX ATC | 7 |
| 2025 | DFlush: DPU-Offloaded Flush for Disaggregated LSM-based Key-Value StoresabstractRapid increase of storage and network bandwidth incurs higher CPU consumption in modern data systems. This phenomenon is particularly evident for log-structured merged key-value stores (LSM-KVS), which rely on resource-intensive background operations to flush and compact disk data. While extensive research has been conducted to reduce the CPU overhead of background compaction, less attention has been paid to background flushing, which can also consume a significant amount of valuable CPU cycles and disrupt CPU caches, ultimately impacting overall performance. In this paper, we propose DFlush, a novel solution that uses DPUs to offload background flush operations to reduce its CPU cost. DPUs are an appealing choice for this goal due to their cost-effectiveness, ease of programming, and widespread deployment. However, their complex hardware architecture requires careful design of both the data and control planes. To fully harness the DPU's capabilities, DFlush decomposes a flush job into fine-grained steps, mapped them to DPU hardware units, and accelerates them through pipeline, data, and channel parallelism, ensuring data-plane efficiency. It also introduces an adaptive control plane that dynamically schedules flush jobs from different LSM-KVS instances based on their priority, reducing write stall and tail latency. Our experiments on a real DPU platform with an industrial-grade LSM-KVS show that DFlush delivers higher throughput, significantly lower tail latency, and saves up to dozens of CPU cores per LSM-KVS server while reducing energy consumption. Chen Ding 0012, Kai Lu 0002, Quanyi Zhang, Zekun Ye, Ting Yao 0001, Daohui Wang, Huatao Wu, Jiguang Wan 0001 |
Proc. ACM Manag. Data | 8 |
| 2025 | NStore: A High-Performance NUMA-Aware Key-Value Store for Hybrid MemoryabstractEmerging persistent memory (PM) promises near-DRAM performance, larger capacity, and data persistence, attracting researchers to design PM-based key-value stores. However, existing PM-based key-value stores lack awareness of the Non-Uniform Memory Access (NUMA) architecture on PM, where accessing PM on remote NUMA sockets is considerably slower than accessing local PM. This NUMA-unawareness results in sub-optimal performance when scaling on NUMA. Although DRAM caching alleviates this issue, existing cache policies ignore the performance disparity between remote and local PM accesses, keeping remote PM access as a performance bottleneck when scaling PM stores on NUMA. Furthermore, creating hot data views in each socket's PM fails to eliminate remote PM writes and, worse, induces additional local PM writes. This paper presents NStore, a high-performance NUMA-aware key-value store for the PM-DRAM hybrid memory. NStore introduces a NUMA-aware cache replacement strategy, called Remote Access First (RAF) cache in DRAM, to minimize remote PM accesses. In addition, NStore deploys Nlog, a write-optimized log-structured persistent storage, purposed to eliminate remote PM writes. NStore further mitigates the NUMA impacts through localized scan operations, efficient garbage collection, and multi-thread recovery for Nlog. Evaluations show that NStore outperforms state-of-the-art PM-based key-value stores, achieving up to 13.9$\times$and 11.2$\times$higher write and read throughput, respectively. Zhonghua Wang 0001, Kai Lu 0002, Jiguang Wan 0001, Hong Jiang 0001, Zeyang Zhao, Biliang Lai, Guokuan Li, Changsheng Xie 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Value-Driven Mixed-Precision Quantization for Patch-Based Inference on MicrocontrollersabstractDeploying neural networks on microcontroller units (MCUs) presents substantial challenges due to their constrained computation and memory resources. Previous researches have explored patch-based inference as a strategy to conserve memory without sacrificing model accuracy. However, this technique suffers from severe redundant computation overhead, leading to a substantial increase in execution latency. A feasible solution to address this issue is mixed-precision quantization, but it faces the challenges of accuracy degradation and a time-consuming search time. In this paper, we propose QuantMCU, a novel patch-based inference method that utilizes value-driven mixed-precision quantization to reduce redundant computation. We first utilize value-driven patch classification (VDPC) to maintain the model accuracy. VDPC classifies patches into two classes based on whether they contain outlier values. For patches containing outlier values, we apply 8-bit quantization to the feature maps on the dataflow branches that follow. In addition, for patches without outlier values, we utilize value-driven quantization search (VDQS) on the feature maps of their following dataflow branches to reduce search time. Specifically, VDQS introduces a novel quantization search metric that takes into account both computation and accuracy, and it employs entropy as an accuracy representation to avoid additional training. VDQS also adopts an iterative approach to determine the bitwidth of each feature map to further accelerate the search process. Experimental results on real-world MCU devices show that QuantMCU can reduce computation by 2.2x on average while maintaining comparable model accuracy compared to the state-of-the-art patch-based inference methods. Shenglin He, Kai Lu 0002, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang, Jing Xiao 0006 |
DATE | 6 |
| 2024 | Enhancing Anomalous Sound Detection with Multi-Level Memory BankabstractAbnormal sound detection (ASD) is crucial for the timely detection of machine faults in industrial scenarios and has emerged as a popular topic. However, exhaustively collecting all ever-changing anomalous samples is impractical for the associated time and cost. Under unsupervised conditions, identifying rare or even unseen abnormal sounds from a large set of normal samples is a notable challenge in the real-world setting. To address this, we propose a novel ASD method based on a multi-level memory bank to estimate the distribution of normal samples in the latent space. We employ a distance-based metric to distinguish inliers from outliers, leveraging high, mid, and low-level features to improve accuracy. We also propose an acoustic-aware farthest embedding sampling algorithm for inference acceleration and memory bank reduction. Experimental results demonstrate our method outperforms existing methods for anomaly detection. Additionally, we analyze the effect of multilevel and acoustic-aware farthest embedding sampling methods, respectively. Baoping Deng, Jinggang Chen, Zhenhou Hong, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
IJCNN | 6 |
| 2024 | PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action RecognitionabstractRecognizing human actions from point cloud sequence has attracted tremendous attention from both academia and industry due to its wide applications. However, most previous studies on point cloud action recognition typically require complex networks to extract intra-frame spatial features and inter-frame temporal features, resulting in an excessive number of redundant computations. This leads to high latency, rendering them impractical for real-world applications. To address this problem, we propose a Plane-Fit Redundancy Encoding point cloud sequence network named PRENet. The primary concept of our approach involves the utilization of plane fitting to mitigate spatial redundancy within the sequence, concurrently encoding the temporal redundancy of the entire sequence to minimize redundant computations. Specifically, our network comprises two principal modules: a Plane-Fit Embedding module and a Spatio-Temporal Consistency Encoding module. The Plane-Fit Embedding module capitalizes on the observation that successive point cloud frames exhibit unique geometric features in physical space, allowing for the reuse of spatially encoded data for temporal stream encoding. The Spatio-Temporal Consistency Encoding module amalgamates the temporal structure of the temporally redundant part with its corresponding spatial arrangement, thereby enhancing recognition accuracy. We have done numerous experiments to verify the effectiveness of our network. The experimental results demonstrate that our method achieves almost identical recognition accuracy while being nearly four times faster than other state-of-the-art methods. Shenglin He, Xiaoyang Qu, Jiguang Wan 0001, Guokuan Li, Jianzong Wang |
IJCNN | 3 |
| 2024 | Gecko: Resource-Efficient and Accurate Queries in Real-Time Video Streams at the EdgeabstractSurveillance cameras are ubiquitous nowadays and users’ increasing needs for accessing real-world information (e.g., finding abandoned luggage) have urged object queries in real-time videos. While recent real-time video query processing systems exhibit excellent performance, they lack utility in deployment in practice as they overlook some crucial aspects, including multi-camera exploration, resource contention, and content awareness. Motivated by these issues, we propose a framework Gecko, to provide resource-efficient and accurate real-time object queries of massive videos on edge devices. Gecko (i) obtains optimal models from the model zoo and assigns them to edge devices for executing current queries, (ii) optimizes resource usage of the edge cluster at runtime by dynamically adjusting the frame query interval of each video stream and forking/joining running models on edge devices, and (iii) improves accuracy in changing video scenes by fine-grained stream transfer and continuous learning of models. Our evaluation with real-world video streams and queries shows that Gecko achieves up to 2x more resource efficiency gains and increases overall query accuracy by at least 12% compared with prior work, further delivering excellent scalability for practical deployment. Liang Wang 0057, Xiaoyang Qu, Jianzong Wang, Guokuan Li, Jiguang Wan 0001, Song Guo 0001, Jing Xiao 0006 |
INFOCOM | 5 |
| 2024 | SepHash: A Write-Optimized Hash Index On Disaggregated Memory via Separate Segment StructureabstractDisaggregated memory separates compute and memory resources into independent pools connected by fast RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. Hash indexes provide high-performance single-point operations and are widely used in distributed systems and databases. However, under disaggregated memory, existing hash indexes suffer from write performance degradation due to high resize overhead and concurrency control overhead. Traditional write-optimized hash indexes are not efficient for disaggregated memory and sacrifice read performance. In this paper, we propose SepHash, a write-optimized hash index for disaggregated memory. First, SepHash proposes a two-level separate segment structure that significantly reduces the bandwidth consumption of resize operations. Second, SepHash employs a low-latency concurrency control strategy to eliminate unnecessary mutual exclusion and check overhead during insert operations. Finally, SepHash designs an efficient cache and filter to accelerate read operations. The evaluation results show that, compared to state-of-the-art distributed hash indexes, SepHash achieves a 3.3X higher write performance while maintaining comparable read performance. Xinhao Min, Kai Lu 0002, Jiguang Wan 0001, Changsheng Xie 0001, Daohui Wang, Ting Yao 0001, Huatao Wu |
Proc. VLDB Endow. | 4 |
| 2024 | D2Comp: Efficient Offload of LSM-tree Compaction with Data Processing Units on Disaggregated StorageabstractLSM-based key-value stores suffer from sub-optimal performance due to their slow and heavy background compactions. The compaction brings severe CPU and network overhead on high-speed disaggregated storage. This article further reveals that data-intensive compression in compaction consumes a significant portion of CPU power. Moreover, the multi-threaded compactions cause substantial CPU contention and network traffic during high-load periods. Based on the above observations, we propose fine-grained dynamical compaction offloading by leveraging the modern Data Processing Unit (DPU) to alleviate the CPU and network overhead. To achieve this, we first customized a file system to enable efficient data access for DPU. We then leverage the Arm cores on the DPU to meet the burst CPU and network requirements to reduce resource contention and data movement. We further employ dedicated hardware-based accelerators on the DPU to speed up the compression in compactions. We integrate our DPU-offloaded compaction with RocksDB and evaluate it with NVIDIA’s latest Bluefield-2 DPU on a real system. The evaluation shows that the DPU is an effective solution to solve the CPU bottleneck and reduce data traffic of compaction. The results show that compaction performance is accelerated by 2.86 to 4.03 times, system write and read throughput is improved by up to 3.2 times and 1.4 times respectively, and host CPU contention and network traffic are effectively reduced compared to the fine-tuned CPU-only baseline. Chen Ding 0012, Jian Zhou 0004, Kai Lu 0002, Sicen Li, Yiqin Xiong, Jiguang Wan 0001 |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated MemoryabstractDisaggregated memory separates compute and memory resources into independent pools connected by RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing RDMA-based distributed transactions on disaggregated memory suffer from severe long-tail latency under high-contention workloads. In this article, we propose Scythe, a novel low-latency RDMA-enabled distributed transaction system for disaggregated memory. Scythe optimizes the latency of high-contention transactions in three approaches: (1) Scythe proposes a hot-aware concurrency control policy that uses optimistic concurrency control (OCC) to improve transaction processing efficiency in low-conflict scenarios. Under high conflicts, Scythe designs a timestamp-ordered OCC (TOCC) strategy based on fair locking to reduce the number of retries and cross-node communication overhead. (2) Scythe presents an RDMA-friendly timestamp service for improved timestamp management. And, (3) Scythe designs an RDMA-optimized RPC framework to improve RDMA bandwidth utilization. The evaluation results show that, compared with state-of-the-art distributed transaction systems, Scythe achieves more than 2.5× lower latency with 1.8× higher throughput under high-contention workloads. Kai Lu 0002, Siqi Zhao, Haikang Shan, Guokuan Li, Jiguang Wan 0001, Ting Yao 0001, Huatao Wu, Daohui Wang |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | WIPE: A Write-Optimized Learned Index for Persistent MemoryabstractLearned Index, which utilizes effective machine learning models to accelerate locating sorted data positions, has gained increasing attention in many big data scenarios. Using efficient learned models, the learned indexes build large nodes and flat structures, thereby greatly improving the performance. However, most of the state-of-the-art learned indexes are designed for DRAM, and there is hence an urgent need to enable high-performance learned indexes for emerging Non-Volatile Memory (NVM). In this article, we first evaluate and analyze the performance of the existing learned indexes on NVM. We discover that these learned indexes encounter severe write amplification and write performance degradation due to the requirements of maintaining large sorted/semi-sorted data nodes. To tackle the problems, we propose a novel three-tiered architecture of write-optimized persistent learned index, which is named WIPE , by adopting unsorted fine-granularity data nodes to achieve high write performance on NVM. Thereinto, we devise a new root node construction algorithm to accelerate searching numerous small data nodes. The algorithm ensures stable flat structure and high read performance in large-size datasets by introducing an intermediate layer (i.e., index nodes) and achieving accurate prediction of index node positions from the root node. Our extensive experiments on Intel DCPMM show that WIPE can improve write throughput and read throughput by up to 3.9× and 7×, respectively, compared to the state-of-the-art learned indexes. Also, WIPE can recover from a system crash in ∼ 18 ms. WIPE is free as an open-source software package. 1 Zhonghua Wang 0001, Chen Ding 0012, Fengguang Song, Kai Lu 0002, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001, Guokuan Li |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | Rcmp: Reconstructing RDMA-Based Memory Disaggregation via CXLabstractMemory disaggregation is a promising architecture for modern datacenters that separates compute and memory resources into independent pools connected by ultra-fast networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing memory disaggregation solutions based on remote direct memory access (RDMA) suffer from high latency and additional overheads including page faults and code refactoring. Emerging cache-coherent interconnects such as CXL offer opportunities to reconstruct high-performance memory disaggregation. However, existing CXL-based approaches have physical distance limitation and cannot be deployed across racks. In this article, we propose Rcmp, a novel low-latency and highly scalable memory disaggregation system based on RDMA and CXL. The significant feature is that Rcmp improves the performance of RDMA-based systems via CXL, and leverages RDMA to overcome CXL’s distance limitation. To address the challenges of the mismatch between RDMA and CXL in terms of granularity, communication, and performance, Rcmp (1) provides a global page-based memory space management and enables fine-grained data access, (2) designs an efficient communication mechanism to avoid communication blocking issues, (3) proposes a hot-page identification and swapping strategy to reduce RDMA communications, and (4) designs an RDMA-optimized RPC framework to accelerate RDMA transfers. We implement a prototype of Rcmp and evaluate its performance by using micro-benchmarks and running a key-value store with YCSB benchmarks. The results show that Rcmp can achieve 5.2× lower latency and 3.8× higher throughput than RDMA-based systems. We also demonstrate that Rcmp can scale well with the increasing number of nodes without compromising performance. Zhonghua Wang 0001, Yixing Guo, Kai Lu 0002, Jiguang Wan 0001, Daohui Wang, Ting Yao 0001, Huatao Wu |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | A Contract-aware and Cost-effective LSM Store for Cloud Storage with Low Latency SpikesabstractCloud storage is gaining popularity because features such as pay-as-you-go significantly reduce storage costs. However, the community has not sufficiently explored its contract model and latency characteristics. As LSM-Tree-based key-value stores (LSM stores) become the building block for numerous cloud applications, how cloud storage would impact the performance of key-value accesses is vital. This study reveals the significant latency variances of Amazon Elastic Block Store (EBS) under various I/O pressures, which challenges LSM store read performance on cloud storage. To reduce the corresponding tail latency, we propose Calcspar, a contract-aware LSM store for cloud storage, which efficiently addresses the challenges by regulating the rate of I/O requests to cloud storage and absorbing surplus I/O requests with the data cache. We specifically developed a fluctuation-aware cache to lower the high latency brought on by workload fluctuations. Additionally, we build a congestion-aware IOPS allocator to reduce the impact of LSM store internal operations on read latency. We evaluated Calcspar on EBS with different real-world workloads and compared it to the cutting-edge LSM stores. The results show that Calcspar can significantly reduce tail latency while maintaining regular read and write performance, keeping the 99 th percentile latency under 550μs and reducing average latency by 66%. In addition, Calcspar has lower write prices and average latency compared to Cloud NoSQL services offered by cloud vendors. Yuanhui Zhou, Jian Zhou 0004, Kai Lu 0002, Shuning Chen, Jiguang Wan 0001 |
ACM Trans. Storage | 9 |
| 2024 | PeakFS: An Ultra-High Performance Parallel File System via Computing-Network-Storage Co-Optimization for HPC ApplicationsabstractEmerging high-performance computing (HPC) applications with diverse workload characteristics impose greater demands on parallel file systems (PFSs). PFSs also require more efficient software designs to fully utilize the performance of modern hardware, such as multi-core CPUs, Remote Direct Memory Access (RDMA), and NVMe SSDs. However, existing PFSs expose great limitations under these requirements due to limited multi-core scalability, unaware of HPC workloads, and disjointed network-storage optimizations. In this article, we present PeakFS, an ultra-high performance parallel file system via computing-network-storage co-optimization for HPC applications. PeakFS designs a shared-nothing scheduling system based on link-reduced task dispatching with lock-free queues to reduce concurrency overhead. Besides, PeakFS improves I/O performance with flexible distribution strategies, memory-efficient indexing, and metadata caching according to HPC I/O characteristics. Finally, PeakFS shortens the critical path of request processing through network-storage co-optimizations. Experimental results show that the metadata and data performance of PeakFS reaches more than 90% of the hardware limits. For metadata throughput, PeakFS achieves a 3.5–19× improvement over GekkoFS and outperforms BeeGFS by three orders of magnitude. Haomai Yang, Kai Lu 0002, Wenlve Huang, Jiguang Wan 0001, Jian Zhou 0004, Fei Wu 0005, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | DoW-KV: A DPU-offloaded and Write-optimized Key-Value Store on Disaggregated Persistent MemoryabstractDisaggregated Persistent Memory (DPM) is a promising technology offering elasticity, high resource utilization, persistent data storage, and lower power consumption. While building KV stores on the DPM benefits from these merits, achieving efficient writes also faces two primary challenges: 1) limited scalability caused by the underused PM bandwidth, and 2) limited CPU resources the persistent memory server (PMS) can provide. Integrating the SmartNIC such as the Data Processing Unit (DPU) into the DPM gives developers the chance to optimize writing to KV stores by utilizing both the memory and processor of DPU. However, simple offloading cannot make full use of the DPU’s potential capacity. To address these challenges, we propose DoW-KV, a persistent hash KV store on DPM. DoW-KV employs a two-tier hash index consisting of a DPU cache table in DPU memory and multiple PM persistent tables on the PM. It relocates small random writes to the DPU memory and consolidates them to the PM at a coarse granularity. Furthermore, DoW-KV uses DPU-offloaded step merge and a coroutine-based asynchronous processing framework to efficiently manage the PM persistent tables. DoW-KV also introduces a client-mixed read strategy to boost key searching on the two-tier hash index. Experimental results show that DoW-KV outperforms the state-of-the-art DINOMO by 2.1× and 1.3× in the Put and Get operations, respectively. Guokuan Li, Jiguang Wan 0001, Junyue Wang, Ting Yao 0001, Huatao Wu, Daohui Wang |
CLUSTER | 3 |
| 2023 | Shoggoth: Towards Efficient Edge-Cloud Collaborative Real-Time Video Inference via Adaptive Online LearningabstractThis paper proposes Shoggoth, an efficient edge-cloud collaborative architecture, for boosting inference performance on real-time video of changing scenes. Shoggoth uses online knowledge distillation to improve the accuracy of models suffering from data drift and offloads the labeling process to the cloud, alleviating constrained resources of edge devices. At the edge, we design adaptive training using small batches to adapt models under limited computing power, and adaptive sampling of training frames for robustness and reducing bandwidth. The evaluations on the realistic dataset show 15%–20% model accuracy improvement compared to the edge-only strategy and fewer network costs than the cloud-only strategy. Liang Wang 0057, Kai Lu 0002, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Guokuan Li, Jing Xiao 0006 |
DAC | 6 |
| 2023 | Detecting Out-of-Distribution Examples Via Class-Conditional Impressions ReappearingabstractOut-of-distribution (OOD) detection aims at enhancing standard deep neural networks to distinguish anomalous inputs from original training data. Previous progress has introduced various approaches where the in-distribution training data and even several OOD examples are prerequisites. However, due to privacy and security, auxiliary data tends to be impractical in a real-world scenario. In this paper, we propose a data-free method without training on natural data, called Class-Conditional Impressions Reappearing (C2IR), which utilizes image impressions from the fixed model to recover class-conditional feature statistics. Based on that, we introduce Integral Probability Metrics to estimate layer-wise class-conditional deviations and obtain layer weights by Measuring Gradient-based Importance (MGI). The experiments verify the effectiveness of our method and indicate that C2IR outperforms other post-hoc methods and reaches comparable performance to the full access (ID and OOD) detection method, especially in the far-OOD dataset (SVHN). Jinggang Chen, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Jing Xiao 0006 |
ICASSP | 5 |
| 2023 | EdgeMA: Model Adaptation System for Real-Time Video Analytics on Edge Devices
Liang Wang 0057, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Guokuan Li, Kaiyu Hu, Guilin Jiang, Jing Xiao 0006 |
ICONIP (1) | 5 |
| 2023 | DComp: Efficient Offload of LSM-tree Compaction with Data Processing UnitsabstractLSM-based Key-value stores suffer from sub-optimal performance due to their slow and heavy background compactions. The compaction overhead shifts to the CPU as the storage performance continuously increases. This paper further reveals that data-intensive compression in compaction consumes a significant portion of CPU power. Moreover, the multi-threaded compactions cause substantial CPU contention during high-load periods. Based on the above observations, we propose fine-grained dynamical compaction offloading by leveraging the modern Data Processing Unit (DPU) to alleviate the CPU overhead. To achieve this, we first employ dedicated hardware-based accelerators on the DPU to speed up the compression in compactions. We then leverage the Arm cores on the DPU to meet the burst CPU requirements to reduce resource contention. We integrate our DPU-offloaded compaction with RocksDB and evaluate it with NVIDIA’s latest Bluefield-2 DPU on a real system. The evaluation shows that the DPU is an effective solution to solve the CPU bottleneck of compaction. The results show that compaction performance is accelerated by 2.86 to 4.03 times, system write and read throughput is improved by up to 3.2 times and 1.4 times respectively, and host CPU contention is effectively reduced compared to the fine-tuned CPU-only baseline. Chen Ding 0012, Jian Zhou 0004, Jiguang Wan 0001, Yiqin Xiong, Sicen Li, Shuning Chen, Kai Lu 0002 |
ICPP | 3 |
| 2023 | GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution DetectionabstractDetecting out-of-distribution (OOD) examples is crucial to guarantee the reliability and safety of deep neural networks in real-world settings. In this paper, we offer an innovative perspective on quantifying the disparities between in-distribution (ID) and OOD data---analyzing the uncertainty that arises when models attempt to explain their predictive decisions. This perspective is motivated by our observation that gradient-based attribution methods encounter challenges in assigning feature importance to OOD data, thereby yielding divergent explanation patterns. Consequently, we investigate how attribution gradients lead to uncertain explanation outcomes and introduce two forms of abnormalities for OOD detection: the zero-deflation abnormality and the channel-wise average abnormality. We then propose GAIA, a simple and effective approach that incorporates Gradient Abnormality Inspection and Aggregation. The effectiveness of GAIA is validated on both commonly utilized (CIFAR) and large-scale (ImageNet-1k) benchmarks. Specifically, GAIA reduces the average FPR95 by 23.10% on CIFAR10 and by 45.41% on CIFAR100 compared to advanced post-hoc methods. Jinggang Chen, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Jing Xiao 0006 |
NeurIPS | 5 |
| 2023 | Calcspar: A Contract-Aware LSM Store for Cloud Storage with Low Latency Spikes
Yuanhui Zhou, Jian Zhou 0004, Shuning Chen, Yanguang Wang, Jiguang Wan 0001 |
USENIX ATC | 9 |
| 2023 | PetaKV: Building Efficient Key-Value Store for File System Metadata on Persistent MemoryabstractPrevious works proposed building file systems and organizing the metadata with KV stores because KV stores handle entries of various sizes efficiently and have excellent scalability. The emergence of the byte-addressable persistent memory (PM) enables metadata service to be faster than before by tailoring the KV store for the PM. However, existing PM-based KV stores cannot handle the workloads of file systems’ metadata well because simply depending on hash tables or trees cannot simultaneously provide fast file accessing and efficient directory traversing. In this paper, we exploit the insight of the metadata operations and propose the PetaKV, a KV store tailored for the metadata management of file systems on PM. PetaKV leverages dual hash indexing to achieve fast file put and get operations. Moreover, it cooperates with PM-tailored peta logs to collocate KV entries for each directory, thus supporting efficient directory scans. Our evaluation indicates PetaKV outperforms state-of-art tree-based KV stores on put, get and scan$2.5\times$,$3.2\times$, and$2.8\times$on average, respectively. Moreover, the file system built with PetaKV achieves$1.2\times$to$6.4\times$speedup compared to those built with tree-based KV stores on the metadata operations. Jian Zhou 0004, Xinhao Min, Jiguang Wan 0001, Ting Yao 0001, Daohui Wang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Accelerating range queries of primary and secondary indices for key-value separationabstractPrimary and secondary indices in LSM-tree-based key-value (KV) stores play significant roles for real-world applications, but they suffer severe I/O amplification due to compaction operations. Prior works show that KV separation can mitigate the I/O amplification under various workloads for either primary or secondary indices. However, range queries of primary and secondary indices only achieve suboptimal efficiency for two reasons: (1) KV separation improves insert/update performance by sacrificing the performance of range queries, (2) range queries of primary and secondary indices may conflict with each other. Chenlei Tang, Jiguang Wan 0001, Zhihu Tan, Guokuan Li |
SoCC | 2 |
| 2022 | RepKV: A Replicated Key-Value Store to Boost Multiple Indices for Key-Value SeparationabstractPrimary and secondary indices are demanded in real-world applications. Recent works show that key-value(KV) separation is efficient to improve queries of multiple secondary indices in LSM-based KV stores. It stores the value in a separate value log and only stores keys and the value address in primary and secondary indices. However, this share-value scheme on multiple indices results in suboptimal efficiency: (1) queries on secondary indices and range queries on all indices cannot fully exploit the bandwidth of SSD devices simultaneously (2) the put operation is inefficient to update secondary indices.To address the above inefficiency, we propose RepKV, a replicated KV store aiming to boost operations of multiple in-dices. Firstly, RepKV uses a primary-backup replication scheme. Each replication stores the same KV pairs but adopts different organizations to exploit the benefit of SSD devices. Secondly, RepKV proposes a lightweight replication scheme to mitigate the extra KV pairs synchronized in replications. Thirdly, RepKV uses a parallel parsing policy to boost the put operation. Experimental results show that RepKV can improve the query performance on secondary indices by up to 22.24%, the range query performance of all indices by up to 31.38%, and the put performance by up to 13.8%. Besides, RepKV can reduce the I/O amplification of replication by up to 3.05x via the lightweight replication scheme. Chenlei Tang, Jiguang Wan 0001, Zhihu Tan, Guokuan Li |
ICCD | 2 |
| 2022 | ADSTS: Automatic Distributed Storage Tuning System Using Deep Reinforcement LearningabstractModern distributed storage systems with the immense number of configurations, unpredictable workloads and difficult performance evaluation pose higher requirements to parameter tuning. Providing an automatic parameter tuning solution for distributed storage systems is in demand. Lots of researches have attempted to build automatic tuning systems based on deep reinforcement learning (RL). However, they have several limitations in the face of these requirements, including lack of parameter spaces processing, less advanced RL models and time-consuming and unstable training process. In this paper, we present and evaluate the ADSTS, which is an automatic distributed storage tuning system based on deep reinforcement learning. A general preprocessing guideline is first proposed to generate standardized tunable parameter domain. Thereinto, Recursive Stratified Sampling without the nonincremental nature is designed to sample huge parameter spaces and Lasso regression is adopted to identify important parameters. Besides, the twin-delayed deep deterministic policy gradient method is utilized to find the optimal values of tunable parameters. Finally, Multi-processing Training and Workload-directed Model Fine-tuning are adopted to accelerate the model convergence. ADSTS is implemented on Park and is used in the real-world system Ceph. The evaluation results show that ADSTS can recommend near-optimal configurations and improve system performance by 1.5 × ∼2.5 × with acceptable overheads. Kai Lu 0002, Guokuan Li, Jiguang Wan 0001, Ruixiang Ma |
ICPP | 3 |
| 2022 | RLRP: High-Efficient Data Placement with Reinforcement Learning for Modern Distributed Storage SystemsabstractModern distributed storage systems with massive data and storage nodes pose higher requirements to the data placement strategy. Furthermore, with emerged new storage devices, heterogeneous storage architecture has become increasingly common and popular. However, traditional strategies expose great limitations in the face of these requirements, especially do not well consider distinct characteristics of heterogeneous storage nodes yet, which will lead to suboptimal performance. In this paper, we present and evaluate the RLRP, a deep reinforcement learning (RL) based replica placement strategy. RLRP constructs placement and migration agents through the Deep-Q-Network (DQN) model to achieve fair distribution and adaptive data migration. Besides, RLRP provides optimal performance for heterogeneous environment by an attentional Long Short-term Memory (LSTM) model. Finally, RLRP adopts Stagewise Training and Model fine-tuning to accelerate the training of RL models with large-scale state and action space. RLRP is implemented on Park and the evaluation results indicate RLRP is a highly efficient data placement strategy for modern distributed storage systems. RLRP can reduce read latency by 10%∼50% in heterogeneous environment compared with existing strategies. In addition, RLRP is used in the real-world system Ceph, which improves the read performance of Ceph by 30%∼40%. Kai Lu 0002, Jiguang Wan 0001, Changhong Fei, Tongliang Deng |
IPDPS | 3 |
| 2022 | Building a Fast and Efficient LSM-tree Store by Integrating Local Storage with Cloud StorageabstractThe explosive growth of modern web-scale applications has made cost-effectiveness a primary design goal for their underlying databases. As a backbone of modern databases, LSM-tree based key–value stores (LSM store) face limited storage options. They are either designed for local storage that is relatively small, expensive, and fast or for cloud storage that offers larger capacities at reduced costs but slower. Designing an LSM store by integrating local storage with cloud storage services is a promising way to balance the cost and performance. However, such design faces challenges such as data reorganization, metadata overhead, and reliability issues. In this article, we propose RocksMash , a fast and efficient LSM store that uses local storage to store frequently accessed data and metadata while using cloud to hold the rest of the data to achieve cost-effectiveness. To improve metadata space-efficiency and read performance, RocksMash uses an LSM-aware persistent cache that stores metadata in a space-efficient way and stores popular data blocks by using compaction-aware layouts. Moreover, RocksMash uses an extended write-ahead log for fast parallel data recovery. We implemented RocksMash by embedding these designs into RocksDB. The evaluation results show that RocksMash improves the performance by up to 1.7 \( \times \) compared to the state-of-the-art schemes and delivers high reliability, cost-effectiveness, and fast recovery. Jiguang Wan 0001, Shuning Chen, Yuanhui Zhou, Hadeel Albahar, Zhihu Tan |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Building GC-free Key-value Store on HM-SMR Drives with ZoneFSabstractHost-managed shingled magnetic recording drives (HM-SMR) are advantageous in capacity to harness the explosive growth of data. For key-value (KV) stores based on log-structured merge trees (LSM-trees), the HM-SMR drive is an ideal solution owning to its capacity, predictable performance, and economical cost. However, building an LSM-tree-based KV store on HM-SMR drives presents severe challenges in maintaining the performance and space utilization efficiency due to the redundant cleaning processes for applications and storage devices (i.e., compaction and garbage collection). To eliminate the overhead of on-disk garbage collection (GC) and improve compaction efficiency, this article presents GearDB , a GC-free KV store tailored for HM-SMR drives. GearDB improves the write performance and space efficiency through three new techniques: a new on-disk data layout, compaction windows, and a novel gear compaction algorithm. We further augment the read performance of GearDB with a new SSTable layout and read ahead mechanism. We implement GearDB with LevelDB, and use zonefs to access a real HM-SMR drive. Our extensive experiments confirm that GearDB achieves both high performance and space efficiency, i.e., on average 1.7× and 1.5× better than LevelDB in random write and read, respectively, with up to 86.9% space efficiency. Ting Yao 0001, Jiguang Wan 0001, Changsheng Xie 0001 |
ACM Trans. Storage | 3 |
| 2022 | TriangleKV: Reducing Write Stalls and Write Amplification in LSM-Tree Based KV Stores With Triangle Container in NVMabstractPopular LSM-tree based key-value stores suffer from suboptimal and unpredictable performance due to write amplification and write stalls that cause application performance to periodically drop to nearly zero. Our preliminary experimental studies reveal that (1) write stalls mainly stem from the significantly large amount of data involved in each compaction between$L_{0}$-$L_{1}$(i.e., the first two levels of LSM-tree), and (2) write amplification increases with the depth of LSM-trees. Existing work mainly focus on reducing write amplification, while only a couple of them target mitigating write stalls. In this paper, we exploit unique features of non-volatile memory (NVM) to address these two limitations and propose TriangleKV, a new LSM-tree based persistent KV store with multi-tier DRAM-NVM-SSD storage. TriangleKV's design principles include performing smaller and cheaper$L_{0}$-$L_{1}$compaction to reduce write stalls while reducing the depth of LSM-trees to mitigate write amplification. To this end, four novel techniques are proposed. First, we relocate and manage the$L_{0}$level in NVM with our proposedtriangle container. Second, the newright-angle side compactionis devised to compact$L_{0}$to$L_{1}$at fine-grained key ranges, thus substantially reducing the amount of compaction data. Third, TriangleKV increases the width of each level to decrease the depth of LSM-trees thus mitigating write amplification. Finally, thecross-row hint searchis introduced for the triangle container to keep adequate read performance. We implement TriangleKV based on MatrixKV and evaluate it on a hybrid DRAM/NVM/SSD system using Intel's latest 3D Xpoint NVM device Optane DC PMM. Evaluation results show that, with the same amount of NVM, TriangleKV outperforms RocksDB, NoveLSM and MatrixKV in 99th-percentile latencies by$5.5\times$,$2.1\times$and$1.1\times$, and random write throughput by$4.9\times$,$3.5\times$and$1.4\times$respectively. Chen Ding 0012, Ting Yao 0001, Hong Jiang 0001, Qiu Cui, Jiguang Wan 0001, Zhihu Tan |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | TridentKV: A Read-Optimized LSM-Tree Based KV Store via Adaptive Indexing and Space-Efficient PartitioningabstractLSM-tree based key-value (KV) stores suffer severe read performance loss due to the leveled structure of the LSM-tree. Especially, when modern storage devices with high bandwidth and low latency are used, the read performance of KV store is seriously affected by inefficient file indexing. Besides, due to the deletion pattern of inserting tombstones, the KV stores based on LSM-tree are faced with the problem of read performance fluctuations that are caused by large-scale data deletion (also referred to as the Read-After-Delete problem). In this article, TridentKV is proposed to improve the read performance of KV stores. An adaptive learned index structure is first designed to speed up file indexing. Also, a space-efficient partition strategy is proposed to solve the Read-After-Delete problem. Besides, asynchronous reading design is adopted, and SPDK is supported for high concurrency and low latency. TridentKV is implemented on RocksDB and the evaluation results indicate that compared with RocksDB, the read performance of TridentKV is improved by 7× to 12× without loss of write performance and TridentKV provides stable read performance even if a large number of deletions or migrations occur. Instead of RocksDB, TridentKV is exploited to store metadata in Ceph, which improves the read performance of Ceph by 20%$\sim$60%. Kai Lu 0002, Jiguang Wan 0001, Changhong Fei, Tongliang Deng |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | FenceKV: Enabling Efficient Range Query for Key-Value SeparationabstractLSM-tree is widely used in key-value stores for big data storage, but it suffers from write amplification brought by frequent compaction operations. An effective solution for this problem is key-value separation, which decouples values from the LSM-tree and stores them in a separate value log. However, existing key-value separation schemes achieve poor range query performance, especially for small key-value pairs, because they focus on mitigating write amplification but neglect access characteristics of the SSD. In this article, we propose FenceKV, which aims to achieve better range query performance while maintaining reasonable update performance for update-intensive workloads. FenceKV employs a new partition method to map values to the storage space based on the key-range to achieve efficient update and range query. Moreover, it adopts a key-range garbage collection policy to mitigate the garbage collection overhead and maintain sequential access for range queries. We compare FenceKV with modern key-value stores with various workloads, and results show that FenceKV can improve the range query performance significantly, while maintaining reasonable update performance compared to the existing designs of key-value separation. Chenlei Tang, Jiguang Wan 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | ComboTree: A Persistent Indexing Structure With Universal Operational Efficiency and ScalabilityabstractTo leverage the larger-than-DRAM capacity and close-to-DRAM performance of persistent memory (PM) to build future memory systems, more scalable and efficient indexing structures are of paramount importance. However, as we evaluate existing PM indexing structures, we find that (1) both ordered and unordered indexing structures cannot support efficiently all KV operation types, i.e., Put, Get, Delete and Scan, and (2) the majority of indexes scale poorly as the PM capacity and dataset increases, i.e., the decreasing throughput of ordered indexes with the growth of dataset and the blocking of foreground requests by costly hash resizing of unordered indexes. To provide better operational efficiency on both point accesses and range queries for persistent memory in the ever-increasing volume of data, this article proposes ComboTree, a three-tiered indexing structure with a sorted key space. In ComboTree, we break the global B+Tree into multiple low height B+Trees (Tier C) and arrange them with a sorted array (Tier B). Further, we accelerate the lookup of the sorted array by cumulative distribution function (Tier A). Last but not least, a background resizing policy is proposed to avoid performance degradation when the capacity of the ComboTree grows. We implement and evaluate ComboTree on Intel’s Optane DCPMM. Test results show that ComboTree delivers$2.1\times -3.6\times$put throughput and$1.5\times -2.1\times$get throughput of the state-of-art sorted indexes. Furthermore, ComboTree is$1.27\times$faster than the efficient B+Tree variant in various scan granularities, and it is open-sourced1. Zhonghua Wang 0001, Ting Yao 0001, Jiguang Wan 0001, Hong Jiang 0001, Qiu Cui |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Building A Fast and Efficient LSM-tree Store by Integrating Local Storage with Cloud StorageabstractThe explosive growth of modern web-scale applications has made cost-effectiveness a primary design goal for their underlying databases. As a backbone of modern databases, LSM-tree based key-value stores (LSM store) face limited storage options. They are either designed for local storage that is relatively small, expensive, and fast or for cloud storage that offers larger capacities at reduced costs but slower. Designing an LSM store by integrating local storage with cloud storage services is a promising way to balance the cost and performance. However, such design faces challenges such as data reorganization and metadata overhead issues. In this paper, we propose ROCKSMASH, a fast and efficient LSM store that uses local storage to store frequently accessed data and metadata while using cloud to hold the rest of the data to achieve cost-effectiveness. To improve metadata space-efficiency and read performance, ROCKSMASH uses an LSM-aware persistent cache that stores metadata in a space-efficient way and stores popular data blocks by using compaction-aware layouts. We implemented ROCKSMASH by embedding these designs into RocksDB. The evaluation results show that ROCKSMASH improves the performance by up to 1.7 × compared to the state-of-the-art schemes and delivers higher reliability and cost-effectiveness. Jiguang Wan 0001, Shuning Chen, Yuanhui Zhou, Hadeel Albahar, Changsheng Xie 0001 |
CLUSTER | 3 |
| 2021 | SW-WAL: Leveraging Address Remapping of SSDs to Achieve Single-Write Write-Ahead LoggingabstractWrite-ahead logging (WAL) has been widely used to provide transactional atomicity in databases, such as SQLite and MySQL/InnoDB. However, the WAL introduces duplicate writes, where changes are recorded in the WAL file and then written to the database file, called checkpointing writes. On the other hand, NAND flash-based SSDs, which have an inherent indirection software layer, called flash translation layer (FTL), become commonplace in modern storage systems. Innovative SSD designs have been proposed to eliminate the WAL overheads by exploiting the FTL, such as providing an atomic write interface or utilizing its address remapping. However, these designs introduce significant performance overheads of maintaining and persisting extra transactional information to guarantee the transactional atomicity or mapping consistency. In this paper, we propose single-write WAL (SW-WAL), a novel cross-layer design, to eliminate WAL-induced duplicate writes on SSDs with minimal overheads. The SSD exposes an address remapping interface to the host, through which the checkpointing writes can be completed without conducting real data writes. To ensure the transactional atomicity and mapping consistency, we make the SSD aware of the transactional writes to the WAL file. Specifically, when transactional data are written to the WAL file, both transactional and mapping semantics are delivered from the host to the SSD and persisted in relevant flash pages as housekeeping metadata without any extra overheads. We implement a prototype of SW-WAL, which runs a popular database SQLite on an emulated NVMe SSD. Experimental results show that SW-WAL improves the database performance by up to 62% compared with original SQLite that bears the WAL overheads and up to 32% compared with the state-of-the-art design that eliminates the WAL overheads. Qiulin Wu, You Zhou 0009, Fei Wu 0005, Jiguang Wan 0001, Changsheng Xie 0001 |
DATE | 6 |
| 2021 | DEPS: Exploiting a Dynamic Error Prechecking Scheme to Improve the Read Performance of SSDabstract3-D NAND flash memory is gradually being widely used in solid state drives (SSDs), leading to increasing storage capacity. However, the read performance of SSD is sacrificed for decoding operations which are executed to guarantee the data reliability. No matter whether the data have bit errors, they will be sent to error correcting code (ECC) engine to decode, introducing a high read delay of SSD. Error prechecking can help to avoid the redundant decoding operations for the error-free data, but it induces extra checking overhead to the error data. Motivated by this, we carry out comprehensive experiments to analyze the distribution of bit errors in 3-D NAND flash memory. The preliminary experimental results show that there are a large number of pages read without errors in the early lifetime of 3-D NAND flash memory. Based on the observations and analyses, we propose a model to estimate the error-free ratio, and utilize it to design a dynamic error prechecking scheme (DEPS) to bypass the decoding operation for the error-free data in 3-D NAND flash memory and improve the read performance of SSD. Furthermore, by dividing a large page into small subpages, DEPS releases more error-free data, which significantly improves the read performance of SSD. Evaluation results from real-world traces demonstrate that by implementing DEPS, the average read performance of SSD is enhanced by 35%-55% with 3-D MLC NAND flash memory. Fei Wu 0005, Meng Zhang 0014, Chengmo Yang, Zhonghai Lu, Jiguang Wan 0001, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Disperse Access Considered Energy Inefficiency in Intel Optane DC Persistent Memory ServersabstractThe Intel Optane DC Persistent Memory Module (AEP), which is the first commercial available Non-Volatile Memory (NVM) product, offers comparable performance with DRAM while providing larger capacities and data persistence. Existing researches that substitute NVM with DRAM or hybridize them are either emulator-based or focused on how to improve the energy efficiency for writes. Unfortunately, the energy efficiency of the real AEP system is less explored. Based on real AEP, we observe that even though eliminating the DRAM-like refresh energy consumptions, AEP consumes significant different energy at different performance levels. Specifically, requests with time intervals (dispersed) underperform in both performance and energy efficiency when compared with the case of requests without time intervals (compact). This disparity and parallelism exploitation potentials motivate us to propose Sprint-AEP, an energy-efficiency-oriented scheduling method for AEP-equipped servers. Sprint-AEP fully activates adequate AEPs to serve most of the requests by deferring the write requests and prefetching the hottest data. The remaining AEPs will stay in idle mode with a low idle power to save energy. Besides, we also utilize the read parallelism to accelerate the sync and prefetching processes. Compared with energy-unaware AEP usages, our experimental results show that Sprint-AEP saves up to 26% energy with little performance degradation. Daping Li, Jiguang Wan 0001, Jun Wang 0001, Jian Zhou 0004, Kai Lu 0002, Fei Wu 0005, Changsheng Xie 0001 |
ICDCS | 2 |
| 2020 | MatrixKV: Reducing Write Stalls and Write Amplification in LSM-tree Based KV Stores with Matrix Container in NVM
Ting Yao 0001, Jiguang Wan 0001, Qiu Cui, Hong Jiang 0001, Changsheng Xie 0001, Xubin He |
USENIX ATC | 3 |
| 2020 | BlockHammer: Improving Flash Reliability by Exploiting Process Variation Aware Proactive Failure Predictionabstractnand flash-based storage devices have gained a lot of popularity in recent years. Unfortunately, flash blocks suffer from limited endurance. For guaranteeing flash reliability, flash manufactures also prescribe a specified number of program and erase (P/E) cycles to define the endurance of flash blocks within the same chip. To extend the service lifetime of a flash-based device, existing works also assume that flash blocks have the same endurance and take P/E-based wear-leveling algorithms which evenly distribute P/E cycle across flash blocks in the controller. However, many studies indicate flash blocks exhibit a wide endurance difference due to the fabrication process. The endurance of flash blocks is limited by the weakest block. Thus, the traditional P/E-based block retirement mechanism makes flash blocks underutilized. To best excavate the endurance of all blocks and improve the reliability of flash devices, we present BlockHammer, a process variation aware proactive failure prediction scheme. BlockHammer takes process variation and blocks similarity into consideration, it consists of a block classifier and a block lifetime predictor. Using machine learning technology, we first establish a block classifier to classify flash blocks into different classes. Based on the classification results, we then establish the block lifetime prediction model for different classes. Flash blocks belonging to the same class are assigned the same model. To verify the effectiveness of BlockHammer, we collect block data from a real nand flash-based testing platform by emulating the true application scenario of nand flash. We compare the predicted value and the tested value, the experimental results show the proposed proactive failure scheme can achieve more than 92% accuracy for flash blocks. Therefore, the block failure point can be accurately predicted using BlockHammer in advance, which greatly enhance the reliability of nand flash. Ruixiang Ma, Fei Wu 0005, Zhonghai Lu, Wenmin Zhong, Qiulin Wu, Jiguang Wan 0001, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Using Error Modes Aware LDPC to Improve Decoding Performance of 3-D TLC NAND Flashabstract3-D triple-level cell (3-D TLC) NAND flash has high storage density and capacity, but degrading data reliability due to high raw bit error rates induced by a certain number of program/erase cycles. To guarantee data reliability, low-density parity-check (LDPC) codes are selected as the error correction codes in modern flash memories because of strong error correction capability. However, directly adopting LDPC codes induces high decoding latency due to iterative updating of log-likelihood ratio (LLR) information in the decoding process. Increasing LLR information accuracy can greatly improve decoding performance. In this paper, we propose EMAL: using error modes aware LDPC codes for further enhancing the decoding performance of 3-D TLC NAND flash. We first obtain 3-D TLC error modes based on an FPGA testing platform, and then exploit the error modes to optimize LLR information and enable the decoding to converge at a high speed. The simulation results show that the decoding performance is significantly improved, resulting in reduced bit error rates and decoding latency. Fei Wu 0005, Meng Zhang 0014, Yajuan Du, Zuo Lu, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | WAS: Wear Aware Superblock Management for Prolonging SSD LifetimeabstractSuperblocks are widely employed in SSDs for improving performance. However, the standard superblock organization which links blocks with the same block ID across planes into one superblock leads to SSDs' ineluctable lifetime waste due to inter-block wear tolerance variations. This work proposes a wear-aware superblock management, called WAS, which (1) dynamically organizes superblocks according to real-time block wear levels to make strong blocks relieve wear on weak ones, and (2) employs a wear-based garbage collection scheme to reduce inter-block wear gap. Comprehensive experiments are carried out in SSDsim. Results show that WAS greatly prolongs SSD lifetime by 51.3% compared with the state-of-the-art superblock management. Shunzhuo Wang, Fei Wu 0005, Chengmo Yang, Jiaona Zhou, Changsheng Xie 0001, Jiguang Wan 0001 |
DAC | 6 |
| 2019 | RAFS: A RAID-Aware File System to Reduce the Parity Update Overhead for SSD RAIDabstractIn a parity-based SSD RAID, small write requests not only accelerate the wear-out of SSDs due to extra writes for updating parities but also deteriorate performance due to associated expensive garbage collection. To mitigate the problem of small writes, a buffer is often added at the RAID controller to absorb overwrites and writes performed to the same stripe. However, this approach achieves only suboptimal efficiency because file layout information is invisible at the block level.This paper proposes RAFS, a RAID-aware file system, which utilizes a RAID-friendly data layout to improve the reliability and performance of SSD-based RAID 5. By leveraging delayed allocation of modern file systems, RAFS employs a stripe-aware buffer policy to coalesce writes to the same file. To reduce parity updates, RAFS compacts buffered updates and flushes back in stripe units to mitigate the parity update overhead. RAFS adopts a stripe-granularity allocation scheme to align writes to stripe boundaries. Experimental results show that RAFS can improve throughput by up to 90%, compared to Ext4. Chenlei Tang, Jiguang Wan 0001, Fei Wu 0005, Changsheng Xie 0001 |
DATE | 2 |
| 2019 | GearDB: A GC-free Key-Value Store on HM-SMR Drives with Gear Compaction
Ting Yao 0001, Jiguang Wan 0001, Ping Huang 0001, Changsheng Xie 0001, Xubin He |
FAST | 2 |
| 2019 | An Active Method to Mitigate the Long Latencies for Host-Aware Shingle Magnetic Recording DrivesabstractShingled Magnetic Recording (SMR) is one of the most promising techniques that satisfy the ever-growing storage volume demands. By overlapping tracks, SMR enormously improves the storage area density, which in turn brings higher storage volumes. However, SMR sacrifices the random write performance for better storage volumes. Current SMR drives propose to remedy this problem by employing an in-drive persistent cache to temporally store incoming writes and migrate them to their disk destinations later on. Unfortunately, cleaning processes for the persistent cache takes up to tens of seconds, and the unpredictable timing of these time-consuming operations chokes normal requests and drastically degrades SMR drive performance. In this paper, we propose to remedy this issue by proactively freeing the persistent cache space so that keeping these lengthy processes transparent with regards to normal requests, therefore reducing the long tails and delivering steady and predictable performance for SMR drive-based storage systems. We prototype our design as a Host-Aware SMR drive aware userspace file system, AM FS, and evaluate it on the real HA-SMR drive with libzbc. Evaluations results show that AM FS reduces the long tails of HA-SMR drives. Jiguang Wan 0001, Ping Huang 0001, Bihua Shu, Chenlei Tang, Changsheng Xie 0001 |
ICPADS | 2 |
| 2019 | SEALDB: An Efficient LSM-tree Based KV Store on SMR Drives with Sets and Dynamic BandsabstractKey-value (KV) stores play an increasingly critical role in supporting diverse large-scale applications in modern data centers hosting terabytes of KV items which even might reside on a single server due to virtualization purposes. The combination of the ever-growing volume of KV items and storage/application consolidation is driving a trend of high storage density for KV stores. Shingled Magnetic Recording (SMR) represents a promising technology for increasing disk capacity, which however comes with the increased complexity of handling random writes. To take the best advantages of SMR drives, applications are expected to work in an SMR-friendly way. In this work, we present SEALDB, a Log-Structured Merge tree (LSM-tree) based key-value store that is specifically optimized for SMR drives via avoiding random writes and the corresponding write amplification on SMR drives. First, for LSM-trees, SEALDB collects and groups participating data of each compaction into sets. Using a set as the basic unit for compactions, SEALDB improves compaction efficiency by reducing random I/Os. Second, SEALDB creates variable sized bands on original HM-SMR drives, named dynamic bands. Dynamic bands store sets in an SMR-friendly way to eliminate the auxiliary write amplification from SMR drives. Third, SEALDB employs two light-weight garbage collection (GC) policies to further improve the space efficiency. We demonstrate the advantages of SEALDB via extensive experiments with various workloads. Overall, SEALDB delivers impressive performance compared with LevelDB, e.g., 3.42×/2.65× faster for random writes (without or with GCs), and 3.96× faster for sequential reads. Ting Yao 0001, Zhihu Tan, Jiguang Wan 0001, Ping Huang 0001, Changsheng Xie 0001, Xubin He |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | A Set-Aware Key-Value Store on Shingled Magnetic Recording Drives with Dynamic BandabstractKey-value (KY) stores play an increasingly critical role in supporting diverse large-scale applications in modern data centers hosting terabytes of KY items which even might reside on a single server due to virtualization purpose. The combination of ever growing volume of KY items and storage/application consolidation is driving a trend of high storage density for KY stores. Shingled Magnetic Recording (SMR) represents a promising technology for increasing disk capacity, but it comes at a cost of poor random write performance and severe I/O amplification. Applications/software working with SMR devices need to be designed and optimized in an SMR-friendly manner. In this work, we present SEALDB, a Log-Structured Merge tree (LSM-tree) based key-value store that is specifically optimized for and works well with SMR drives via adequately addressing the poor random writes and severe I/O amplification issues. First, for LSM-trees, SEALDB concatenates SSTables of each compaction, and groups them into sets. Taking sets as the basic unit for compactions, SEALDB improves compaction efficiency by mitigating random I/Os. Second, SEALDB creates varying size bands on HM-SMR drives, named dynamic bands. Dynamic bands not only accommodate the storage of sets, but also eliminate the auxiliary write amplification from SMR drives. We demonstrate the advantages of SEALDB via extensive experiments in various workloads. Overall, SEALDB delivers impressive performance improvement. Compared with LevelDB, SEALDB is 3.42× faster on random load due to improved compaction efficiency and eliminated auxiliary write amplification on SMR drives. Ting Yao 0001, Zhihu Tan, Jiguang Wan 0001, Ping Huang 0001, Changsheng Xie 0001, Xubin He |
IPDPS | 3 |
| 2018 | Chameleon: An Adaptive Wear Balancer for Flash ClustersabstractNAND flash-based Solid State Devices (SSDs) offer the desirable features of high performance, energy efficiency, and fast growing capacity. Thus, the use of SSDs is increasing in distributed storage systems. A key obstacle in this context is that the natural unbalance in distributed I/O workloads can result in wear imbalance across the SSDs in a distributed setting. This, in turn can have significant impact on the reliability, performance, and lifetime of the storage deployment. Extant load balancers for storage systems do not consider SSD wear imbalance when placing data, as the main design goal of such balancers is to extract higher performance. Consequently, data migration is the only common technique for tackling wear imbalance, where existing data is moved from highly loaded servers to the least loaded ones. In this paper, we explore an innovative holistic approach, Chameleon, that employs data redundancy techniques such as replication and erasure-coding, coupled with endurance-aware write offloading, to mitigate wear level imbalance in distributed SSD-based storage. Chameleon aims to balance the wear among different flash servers while meeting desirable objectives of: extending life of flash servers; improving I/O performance; and avoiding bottlenecks. Evaluation with a 50 node SSD cluster shows that Chameleon reduces the wear distribution deviation by 81% while improving the write performance by up to 33%. Ali Anwar 0001, Yue Cheng 0001, Mohammed Salman, Daping Li, Jiguang Wan 0001, Changsheng Xie 0001, Xubin He, Feiyi Wang, Ali Raza Butt |
IPDPS | 6 |
| 2018 | Workload Scheduling for Massive Storage Systems with Arbitrary Renewable SupplyabstractAs datacenters grow in scale, increasing energy costs and carbon emissions have led data centers to seek renewable energy, such as wind and solar energy. However, tackling the challenges associated with the intermittency and variability of renewable energy is difficult. This paper proposes a scheme called GreenMatch, which deploys an SSD cache to match green energy supplies with a time-shifting workload schedule while maintaining low latency for online data-intensive services. With the SSD cache, the process for a latency-sensitive request to access a disk is divided into two stages: a low-energy/low-latency online stage and a high-energy/high-latency off-line stage. As the process in the latter stage is off-line, it offers opportunities for time-shifting workload scheduling in response to variations of green energy supplies. We also allocate an HDD cache to guarantee data availability when renewable energy is inadequate. Furthermore, we design a novel replacement policy called Inactive P-disk First for the HDD cache to avoid inactive disk accesses. The experimental results show that GreenMatch can make full use of renewable energy while minimizing the negative impacts of intermittency and variability on performance and availability. Daping Li, Xiaoyang Qu, Jiguang Wan 0001, Jun Wang 0001, Xiaozhao Zhuang, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | CooECC: A Cooperative Error Correction Scheme to Reduce LDPC Decoding Latency in NAND FlashabstractThe storage capacity of NAND Flash has increased by scaling down to smaller cell size and using multi-level storage technology, but data reliability is degraded by severer retention errors. To ensure data reliability, error correction codes (ECC) are adopted, such as BCH and low-density parity check (LDPC) codes. However, BCH codes are insufficient when raw bit error rates (RBER) caused by retention errors are high. As a result, BCH codes are inevitably replaced with LDPC codes with stronger error correction capability. Traditional LDPC codes are used to independently correct bit errors in the LSB and MSB pages. Unfortunately, decoding latency in such two pages is significantly unbalanced, MSB pages take much higher latency due to higher RBER, leading to suboptimal flash read performance. This paper proposes a cooperative error correction scheme, called CooECC, to reduce LDPC decoding latency of the MSB page in NAND Flash. By exploiting data error characteristics introduced by retention errors, CooECC integrates the decoding result of the LSB page into the initial information of LDPC decoding for the MSB page, making it more accurate. This in turn enables decoding to converge at a higher rate. Simulation results show that for LDPC schemes with information lengths of 2KB and 4KB, the decoding latency can be reduced by up to 87% and 84%, respectively, when RBER is as high as 8.0 × 10^-3. Meng Zhang 0014, Fei Wu 0005, Yajuan Du, Chengmo Yang, Changsheng Xie 0001, Jiguang Wan 0001 |
ICCD | 6 |
| 2017 | OptiMatch: Enabling an Optimal Match between Green Power and Various Workloads for Renewable-Energy Powered Storage SystemsabstractTo reduce energy consumption and carbon emission, many data centers have deployed (or anticipate to build) their own renewable-energy power plants. However, the renewable energy (such as wind, tide, and solar energy) has the serious issues of intermittency and variability that prevent the green energy from being utilized effectively in practice. To cope with the issues, new power-supply management policies and workload scheduling algorithms have been designed. However, most existing work focuses on power optimization on computation only. In this paper, we introduce a novel scheme called OptiMatch to optimize the match between the power supply and the user-workload demand for massive storage systems that are mostly powered by renewable energy sources. OptiMatch has a hierarchical architecture, which consists of a number of heterogeneous storage devices. OptiMatch systematically utilizes the performance disparities between heterogeneous storage devices (i.e., performance per watt, IOPS/watt) to split the process for every write request into two stages: an on-line stage and a deferred off-line stage. The deferred off-line requests are used to match the green energy supplies. To maximize green energy utilization and minimize power budget without sacrificing quality of service, the fundamental methodology is to make the aggregate power supplies be proportional to the I/O workload demand at any time. To this end, our OptiMatch employs novel co-design optimizations. (1) We propose a dual-drive power control approach that makes the number of active nodes proportional to the workload demand when the green power supply is insufficient, meanwhile be proportional to the green power supply when green power is sufficient. (2) During periods of insufficient green supplies, we exploit virtualization consolidation schemes which enable a fine-grained power control to minimize the grid budgets. (3) During the periods of sufficient green supplies, we design an intelligent workload scheduling scheme which enables a near-optimal off-line requests assignment to maximize the green utilization. The experimental results demonstrate that the new OptiMatch framework can achieve high green utilization (up to 94.9%) with a minor performance degradation (less than 9.8%). Xiaoyang Qu, Jiguang Wan 0001, Fengguang Song, Xiaozhao Zhuang, Fei Wu 0005, Changsheng Xie 0001 |
ICPP | 2 |
| 2017 | DEFT-Cache: A Cost-Effective and Highly Reliable SSD Cache for RAID StorageabstractThis paper proposes a new SSD cache architecture, DEFT-cache, Delayed Erasing and Fast Taping, that maximizes I/O performance and reliability of RAID storage. First of all, DEFT-Cache exploits the inherent physical properties of flash memory SSD by making use of old data that have been overwritten but still in existence in SSD to minimize small write penalty of RAID5/6. As data pages being overwritten in SSD, old data pages are invalidated and become candidates for erasure and garbage collections. Our idea is to selectively delay the erasure of the pages and let these otherwise useless old data in SSD contribute to I/O performance for parity computations upon write I/Os. Secondly, DEFT-Cache provides inexpensive redundancy to the SSD cache by having one physical SSD and one virtual SSD as a mirror cache. The virtual SSD is implemented on HDD but using log-structured data layout, i.e. write data are quickly logged to HDD using sequential write. The dual and redundant caches provide a cost-effective and highly reliable write-back SSD cache. We have implemented DEFT-Cache on Linux system. Extensive experiments have been carried out to evaluate the potential benefits of our new techniques. Experimental results on SPC and Microsoft traces have shown that DEFT-Cache improves I/O performance by 26.81% to 56.26% in terms of average user response time. The virtual SSD mirror cache can absorb write I/Os as fast as physical SSD providing the same reliability as two physical SSD caches without noticeable performance loss. Jiguang Wan 0001, Qing Yang 0001, Xiaoyang Qu, Changsheng Xie 0001 |
IPDPS | 1 |
| 2017 | Exploiting Virtual Metadata Servers to Provide Multi-Level Consistency for Key-Value Object-Based Data StoreabstractDistributed data store is a fundamental building block for various Internet services. For large-scale distributed data store, the scalability and consistency of metadata services are prone to be the bottleneck. Various schemes are proposed to tackle the challenge of scalability and consistency within metadata services. While centralized single-node metadata services with low scalability provide low- overhead consistency maintenance, distributed metadata servers with high scalability often suffer complicated management and high-overhead consistency maintenance. As some key-value object-based storage systems locate and access an object by hashing function (e.g., consistent hashing table), there are no dedicated physical servers for metadata services. For key-value store without dedicated metadata servers, we exploited a scheme called virtual metadata servers (virtual MDS), which can create an opportunity to provide high performance and multi- level consistency. While conventional key-value data store distributes metadata across data nodes, our scheme uses proxy nodes, where virtual disks created, as virtual MDS to hold the metadata of virtual disks. Meanwhile, we also combine the characteristic of virtual disks and metadata services to implement a multi-level consistency strategy for the key-value object-based store without dedicated physical metadata servers. With virtual MDS, we use version information to update data asynchronously and check the version consistency periodically, then correct the stale entries properly. In this way, our virtual MDS can provide multi-level of consistency to cope with different read performance demand from users. The experiment results demonstrate that our scheme with relaxed consistency can enhance random write performance by 50% and improve random read performance by 16% compared with the standard storage system with strict consistency. Xiaozhao Zhuang, Xiaoyang Qu, Zhiyong Lu, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 4 |
| 2017 | DROP: A New RAID Architecture for Enhancing Shared RAID PerformanceabstractEnterprise storage systems are generally shared by multiple servers in a storage area network environment. Our experiments as well as industry reports have shown that disk arrays show poor performance when multiple servers share one RAID due to resource contention as well as frequent disk head movements. We have studied IO performance characteristics of several shared storage settings of practical business operations. To avoid the IO contention, we propose a new dynamic data relocation technique on shared RAID storages, referred to as DROP, dynamic data relocation to optimize performance. DROP allocates/manages a group of cache data areas and relocates/drops the portion of hot data at a predefined sub-array that is a physical partition on the top of the entire shared array. By analyzing the profiling data, we are able to determine the optimal data relocation and partition of disks in the RAID to maximize large sequential block accesses on individual disks and at the same time maximize parallel accesses across disks in the array. As a result, DROP minimizes disk head movements in the array at run time giving rise to fast IO response time. A prototype DROP has been implemented as a software module at the storage target controller. Extensive experiments have been carried out using real world IO workloads to evaluate the performance of the DROP implementation. Experimental results have shown that DROP improves the shared IO performance greatly. The performance improvements in terms of the average IO response time range from 42.06% to 58.34% at no additional hardware cost. Jiguang Wan 0001, Changsheng Xie 0001 |
Comput. J. | 2 |
| 2017 | A reliable and energy-efficient storage system with erasure coding cacheabstractIn modern energy-saving replication storage systems, a primary group of disks is always powered up to serve incoming requests while other disks are often spun down to save energy during slack periods. However, since new writes cannot be immediately synchronized into all disks, system reliability is degraded. In this paper, we develop a high-reliability and energy-efficient replication storage system, named RERAID, based on RAID10. RERAID employs part of the free space in the primary disk group and uses erasure coding to construct a code cache at the front end to absorb new writes. Since code cache supports failure recovery of two or more disks by using erasure coding, RERAID guarantees a reliability comparable with that of the RAID10 storage system. In addition, we develop an algorithm, called erasure coding write (ECW), to buffer many small random writes into a few large writes, which are then written to the code cache in a parallel fashion sequentially to improve the write performance. Experimental results show that RERAID significantly improves write performance and saves more energy than existing solutions. Jiguang Wan 0001, Daping Li, Xiaoyang Qu, Jun Wang 0001, Changsheng Xie 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2017 | A Program Interference Error Aware LDPC Scheme for Improving NAND Flash Decoding PerformanceabstractBy scaling down to smaller cell size, NAND flash has significantly increased the storage capacity in order to lower the unit cost down. However, the reliability is sacrificed due to much higher raw bit error rates. As a result, conventional error correction codes (ECCs), such as BCH codes, are not sufficient. Low-density parity check (LDPC) codes with stronger error correction capability are adopted in NAND flash to guarantee data reliability. However, read performance using LDPC is poor because of its decoding complexity. It has been found that flash cells with fewer electrons are more prone to program interference errors. As a result, program interference errors show the characteristic of value dependence. This characteristic can be exploited and translated into extra information facilitating the decoding convergence. Motivated by this observation, we propose PEAL: a flash program interference error aware LDPC scheme to enhance the decoding performance. PEAL integrates the obtained extra information from the value dependence into the soft-to-hard decision process in LDPC decoding to decrease decoding iterations and improve the decoding convergence speed. Simulation results show that decoding iterations are reduced by up to 69.37% and the decoding convergence speed is improved by up to 2.5×, compared with the normalized min-sum (NMS) algorithm with 2KB information lengths at an approximate raw bit error rate of 11.5 × 10 −3 . Fei Wu 0005, Meng Zhang 0014, Yajuan Du, Xubin He, Ping Huang 0001, Changsheng Xie 0001, Jiguang Wan 0001 |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2017 | Building Efficient Key-Value Stores via a Lightweight Compaction TreeabstractLog-Structure Merge tree (LSM-tree) has been one of the mainstream indexes in key-value systems supporting a variety of write-intensive Internet applications in today’s data centers. However, the performance of LSM-tree is seriously hampered by constantly occurring compaction procedures, which incur significant write amplification and degrade the write throughput. To alleviate the performance degradation caused by compactions, we introduce a lightweight compaction tree (LWC-tree), a variant of LSM-tree index optimized for minimizing the write amplification and maximizing the system throughput. The lightweight compaction drastically decreases write amplification by appending data in a table and only merging the metadata that have much smaller size. Using our proposed LWC-tree, we have implemented three key-value LWC-stores on different storage mediums including Shingled Magnetic Recording (SMR) drives, Solid State Drives (SSD), and conventional Hard Disk Drives (HDDs). The LWC-store is particularly optimized for SMR drives, as it eliminates the multiplicative I/O amplification from both LSM-trees and SMR drives. Due to the lightweight compaction procedure, LWC-store reduces the write amplification by a factor of up to 5× compared to the popular LevelDB key-value store. Moreover, the random write throughput of the LWC-tree on SMR drives is significantly improved by up to 467% even compared with LevelDB on conventional HDDs. Furthermore, LWC-tree has wide applicability and delivers impressive performance improvement in various conditions, including different storage mediums (i.e., SMR, HDD, SSD) and various value sizes and access patterns (i.e., uniform and Zipfian). Ting Yao 0001, Jiguang Wan 0001, Ping Huang 0001, Xubin He, Fei Wu 0005, Changsheng Xie 0001 |
ACM Trans. Storage | 2 |
| 2016 | GreenMatch: Renewable-Aware Workload Scheduling for Massive Storage SystemsabstractAs datacenters grow in scale, increasing energy costs and carbon emissions have led data centers to seek renewable energy, such as wind and solar energy. However, tackling the challenges associated with the intermittent nature and variability of renewable energy is substantial. This paper proposes a scheme called GreenMatch, which deploys an SSD-cache to match green energy supplies with a time-shifting workload schedule while maintaining low latency for online data-intensive services. With the SSD-cache, the process for a latency-sensitive request to access a disk is divided into two stages: a low-energy low-latency online stage and a high-energy high-latency off-line stage. As the process in the latter stage is off-line, it offers opportunities for time-shifting workload scheduling in response to variations of green energy supplies. We also allocate an HDD-cache to guarantee data availability when renewable energy is non-adequate. Furthermore, we design a novel replacement policy called Inactive Disk First for the HDD-cache to avoid inactive disk accesses. The experimental results show that GreenMatch can make full use of renewable energy while minimizing the negative impact of intermittency and variability on performance and availability. Xiaoyang Qu, Jiguang Wan 0001, Jun Wang 0001, Liqiong Liu, Changsheng Xie 0001 |
IPDPS | 2 |
| 2016 | CircularCache: Scalable and Adaptive Cache Management for Massive Storage SystemsabstractIn order to enhance the performance of HDD-based storage systems, low-latency and high-IOPS SSDs are usually deployed as a cache above HDDs. With explosive data growth, a large-scale SSD-based cache tend to adopt partition management for overall cached data distribution across multiple cache nodes. We proposed an adaptive and scalable SSD- based cache called CircularCache, which distributes hot data across multiple cache nodes. The hotter virtual disks deserve more allocated free space in the SSD-cache. This paper exploited a dynamic replacement algorithm called VBQ(VDI-Based Queues) to manage the SSD-cache. The VBQ scheme manages the SSD-cache by dynamically manipulating the upper- bounds and lower-bounds of multiple queues based on the total access number of virtual disks. To mitigate negative impacts of destaging on overall storage performance, the dirty data in the cache will be written back to data nodes during idle time. At the same time, we utilize the redundant storage space in the data nodes as logging area to retain reliability of the dirty data on the SSDcache. The prototype of CircularCache is implemented based on Sheepdog. Experimental results show that CircularCache offers a performance improvement by up to 270% compared with the standard distributed storage system without an SSD-based cache. Liqiong Liu, Xiaoyang Qu, Yubiao Zhang, Xiaodong Yi 0003, Siwang Zeng, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 6 |
| 2016 | DVS: Dynamic Variable-Width Striping RAID for Shingled Write DisksabstractDisk data density improvement will eventually be limited by the super-paramagnetic effect for perpendicular magnetic recording. Of the various new technologies being explored, Shingled Magnetic Recording (SMR) exposes as the most promising one to achieve high areal density and only make little changes to the manufacturing process. At present, high-capacity SMR drives are available from Seagate and HGST. Since SMR is leading next generation disk technology and increasing SMR drives will be used in storage systems, there is a great need to look over the current RAID storage techniques based on HDDs again. In this paper, we proposed a dynamic variable-width striping RAID (DVS-RAID) for SMR drives to reduce the parity updating cost. DVS-RAID never overwrites the old data, but always constructs a new full or partial stripe (variable-width stripe), and writes to the SMR drives through appending. In addition, taking the access characteristics of SMR drives into consideration, we present a new write cache management that exploits both spatial and temporal localities. The experiment with six real-world traces demonstrates that DVS- RAID exhibits a slightly lower performance than HDD- based RAID on update intensive workloads. However the performance of DVS-RAID is better than HDD-based RAID with sequential access, read-dominated workloads or workloads with rarely update. Ting Yao 0001, Xiaoyang Qu, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 4 |
| 2016 | A reliable power management scheme for consistent hashing based distributed key value storage systemsabstractDistributed key value storage systems are among the most important types of distributed storage systems currently deployed in data centers. Nowadays, enterprise data centers are facing growing pressure in reducing their power consumption. In this paper, we propose GreenCHT, a reliable power management scheme for consistent hashing based distributed key value storage systems. It consists of a multi-tier replication scheme, a reliable distributed log store, and a predictive power mode scheduler (PMS). Instead of randomly placing replicas of each object on a number of nodes in the consistent hash ring, we arrange the replicas of objects on nonoverlapping tiers of nodes in the ring. This allows the system to fall in various power modes by powering down subsets of servers while not violating data availability. The predictive PMS predicts workloads and adapts to load fluctuation. It cooperates with the multi-tier replication strategy to provide power proportionality for the system. To ensure that the reliability of the system is maintained when replicas are powered down, we distribute the writes to standby replicas to active servers, which ensures failure tolerance of the system. GreenCHT is implemented based on Sheepdog, a distributed key value storage system that uses consistent hashing as an underlying distributed hash table. By replaying 12 typical real workload traces collected from Microsoft, the evaluation results show that GreenCHT can provide significant power savings while maintaining a desired performance. We observe that GreenCHT can reduce power consumption by up to 35%–61%. Jiguang Wan 0001, Jun Wang 0001, Changsheng Xie 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2016 | H-Scale: A Fast Approach to Scale Disk Arrays via Hybrid Stripe DeploymentabstractTo satisfy the explosive growth of data in large-scale data centers, where redundant arrays of independent disks (RAIDs), especially RAID-5, are widely deployed, effective storage scaling and disk expansion methods are desired. However, a way to reduce the data migration overhead and maintain the reliability of the original RAID are major concerns of storage scaling. To address these problems, we propose a new RAID scaling scheme, H-Scale, to achieve fast RAID scaling via hybrid stripe layouts. H-Scale takes advantage of the loose restriction of stripe structures to choose migrated data and to create hybrid stripe structures. The main advantages of our scheme include: (1) dramatically reducing the data migration overhead and thus speeding up the scaling process, (2) maintaining the original RAID’s reliability, (3) balancing the workload among disks after scaling, and (4) providing a general scaling approach for different RAID levels. Our theoretical analysis show that H-Scale outperforms existing scaling solutions in terms of data migration, I/O overheads, and parity update operations. Evaluation results on a prototype implementation demonstrate that H-Scale speeds up the online scaling process by up to 60% under SPC traces, and similar improvements on scaling time and user response time are also achieved by evaluations using standard benchmarks. Jiguang Wan 0001, Xubin He, Junyao Li, Changsheng Xie 0001 |
ACM Trans. Storage | 1 |
| 2016 | Design and Implementation of a Hybrid Shingled Write Disk SystemabstractDisk data density improvement will eventually be limited by the super-paramagnetic effect for perpendicular recording. While various approaches to this problem have been proposed, Shingled Magnetic Recording (SMR) holds great promise to mitigate the problem of density scaling cost-effectively by overlapping data tracks. However, the inherent properties of SMR limit Shingled Write Disk (SWD) applicability since writing data to one track destroys the data previously-stored on the overlapping tracks. As a result, various data layout management designs have been proposed. In this paper, we present a hybrid wave-like shingled recording (HWSR) disk system, which can improve both the performance and the capacity of a shingled write disk. We propose a novel segment-based data layout management and a new wave-like shingled recording that overlaps adjacent tracks from two opposite radial directions. This new scheme can not only efficiently reduce the write amplification, but also double the areal density of conventional circular log-based shingled recording. A new replacement policy based on least write amplification is also devised to manage the hybrid system to effectively eliminate the performance degradation. Our measurements on HWSR implemented in Linux kernel 2.6.35.6 show that it provides superb performance. For example, HWSR reduces the average I/O response time by an order of magnitude compared to S-block forFinancial1trace, and provides up to 3.7 speedup over standard hard disks without using shingled magnetic recording technology. Jiguang Wan 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | GreenCHT: A power-proportional replication scheme for consistent hashing based key value storage systemsabstractDistributed key value storage systems are widely used by many popular networking corporations. Nevertheless, server power consumption has become a growing concern for key value storage system designers since the power consumption of servers contributes substantially to a data center's power bills. In this paper, we propose GreenCHT, a power-proportional replication scheme for consistent hashing based key value storage systems. GreenCHT consists of a power-aware replication strategy — multi-tier replication strategy and a centralized power control service — predictive power-mode scheduler. The multitier replication provides power-proportionality and ensures data availability, reliability, consistency, as well as fault-tolerance of the whole system. The predictive power-mode scheduler component predicts workloads and exploits load fluctuation to schedule nodes to be powered-up and powered-down. GreenCHT is implemented based on Sheepdog, a distributed key value system that uses consistent hashing as an underlying distributed hash table. By replicating twelve real workload traces collected from Microsoft, the evaluation results show that GreenCHT can provide significant power savings while maintaining an acceptable performance. We observed that GreenCHT can reduce power consumption by up to 35%–61%. Jiguang Wan 0001, Jun Wang 0001, Changsheng Xie 0001 |
MSST | 2 |
| 2015 | ThinRAID: Thinning Down RAID Array for Energy ConservationabstractThe current power managements in RAID array are mostly designed to conserve energy by spinning down partial disks of standard RAID architecture. However, spinning down several disks not only decreases disk parallelism, but also creates new problems, for example, partial chunks of the stripe cannot be accessed directly or multiple chunks of the same stripe are stored on the same disk, which affect spatial locality. We refer these problems as stripe degradation, which results in further performance degradation. To avoid such problems, this paper proposes a new RAID storage architecture called ThinRAID, which uses a subset of disks to build a capacity-adaptive RAID array based on the volume of the data set. Also, the other non-essential disks are spun down to save energy. When the workload is projected to become heavier based on our forecast model, data are migrated to disks that have recently transitioned from standby to active. Furthermore, we also propose a novel data reorganization algorithm that can minimize data migration. We have implemented ThinRAID in the Linux kernel and evaluated its performance and energy efficiency by replaying seven representative traces. Experimental results show that ThinRAID can save 15-27 percent on energy on average over conventional RAID, with minimum performance degradation. In comparison to PARAID, ThinRAID achieves up to 62 percent performance improvement. Jiguang Wan 0001, Xiaoyang Qu, Jun Wang 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | A New Parity-Based Migration Method to Expand RAID-5abstractTo expand the capacity of a RAID-5 array with additional disks, data have to be migrated between disks to leverage extra space and performance gain. Conventional methods for expanding RAID-5 are very slow because they have to migrate almost all existing data and recalculate all parity blocks. This paper proposes a new online expansion method for RAID-5, named parity-based migration (PBM). This method only migrates blocks that form a special parallelogram with one side consisting of only parity blocks. When adding m disks to a RAID-5 with n disks, PBM achieves the minimal data migration which only needs to move m/(n+m) of all data blocks. Furthermore, no parity blocks are recalculated during the expansion. After expansion, although the RAID is not a standard RAID-5 distribution, the parity blocks are distributed evenly. Experimental results based on extensive trace-driven show that, on average, PBM can reduce the time of expansion by 73.6 percent while only reduces the performance of the expanded RAID by 1.83 percent when compared with Multiple-Device (MD), a toolkit provided in Linux kernel. Jiguang Wan 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | ${\rm S}^{2}$-RAID: Parallel RAID Architecture for Fast Data RecoveryabstractAs disk volume grows rapidly with terabyte disk becoming a norm, RAID reconstruction process in case of a failure takes prohibitively long time. This paper presents a new RAID architecture, S2-RAID, allowing the disk array to reconstruct very quickly in case of a disk failure. The idea is to form skewed sub-arrays in the RAID structure so that reconstruction can be done in parallel dramatically speeding up data reconstruction process and hence minimizing the chance of data loss. We analyse the data recovery ability of this architecture and show its good scalability. A prototype S2-RAID system has been built and implemented in the Linux operating system for the purpose of evaluating its performance potential. Real world I/O traces including SPC, Microsoft, and a collection of a production environment have been used to measure the performance of S2-RAID as compared to existing baseline software RAID5, Parity Declustering, and RAID50. Experimental results show that our new S2-RAID speeds up data reconstruction time by a factor 2 to 4 compared to the traditional RAID. Meanwhile, S2-RAID keeps comparable production performance to that of the baseline RAID layouts while online RAID reconstruction is in progress. Jiguang Wan 0001, Changsheng Xie 0001, Qing Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Robot: An efficient model for big data storage systems based on erasure codingabstractIt is well-known that with the explosive growth of data, the age of big data has arrived. How to save huge amounts of data is of great importance to both industry and academia. This paper puts forward a solution based on coding technologies in big data system that store a lot of cold data. By studying existing coding technologies and big data systems, we can not only maintain the system's reliability, but also improve the security and the utilization of storage systems. Due to the remarkable reliability and space saving rate of coding technologies, importing coding schema in to big data systems becomes prerequisite. In our presented schema, the storage node is divided into several virtual nodes to keep load balancing. By setting up different virtual node storage groups for different codec server, we can ensure system availability. And by utilizing the parallel decoding computing of the node and the block of data, we can also reduce the system recovery time when data is corrupted. Additionally, different users set different coding parameters can improve the robustness of big data storage systems. We configure various data block m and calibration block k to improve the utilization rate in the quantitative experiments. The results shows that parallel decoding speed can rise up two times than the past serial decoding speed. The encoding efficiency with ICRS coding is 34.2% higher than using CRS and 56.5% more than using RS coding equally. The decoding rate by using ICRS is 18.1% higher than using CRS and 31.1% higher than using RS averagely. Jianzong Wang, Jiguang Wan 0001, Changlin Long, Wenjuan Bi |
IEEE BigData | 4 |
| 2013 | A reliability optimization method for RAID-structured storage systems based on active data migration
Zhihu Tan, Jiguang Wan 0001, Changsheng Xie 0001 |
J. Syst. Softw. | 3 |
| 2012 | RO-BURST: A Robust Virtualization Cost Model for Workload Consolidation over CloudsabstractAs more public cloud computing platforms are emerging in the market, a great challenge for these Infrastructure as a Server (IaaS) providers is how to measure the cost and charge the Software as a Service (SaaS) clients for the cloud computing services. This problem is compounded as virtualization technology is deployed in many cloud platforms to consolidate servers and improve their utilization. This paper studies three different but related models for apportioning costs in a private or public cloud environment supported by virtualized data centers. With given workload placement scenarios and randomly selected workloads, these models estimate the cost for each workload. Through simulations and thorough comparisons of the results, we finally champion the RO-BURST model tailored for the service providers' need, that is characterized by robustness and burstiness. What is more, we import Cost Volatility Factors to ensure that our model is able to adjust itself to the market and multiform demands in power and hardware components, such as disks and CPU, showing its compatibility and extensibility. We also come up with a pricing strategy with respect to servers the workload employs, which generates an applicable and less placement-sensitive fee for the clients. Jianzong Wang, Jiguang Wan 0001 |
CCGRID | 4 |
| 2012 | High Performance and High Capacity Hybrid Shingled-Recording Disk SystemabstractAreal density scaling in magnetic hard drives isin jeopardy as magnetic particles become unstable when they are sufficiently small. Shingled recording holds great promise to mitigate the problem of density scaling cost-effectively by overlapping data tracks. However, this innovative technology suffers severely from slow small writes. This prevents shingle recording from being widely adapted in practice. This paper presents a new hybrid storage architecture that combines a shingled-recording magnetic disk and a fast SSD cache to achieve a high-capacity storage system without any compromise to performance. We propose a new wave-like shingled recording that overlaps adjacent tracks from two opposite radial directions. This new schemes doubles the areal density of conventional circular log-based shingled recording. We also design a new replacement strategy to manage the hybrid system to effectively eliminate the performance degradation. We evaluate our design based on a prototype implementation. Experimental results under 12 I/O workloads show that our hybrid system exhibits a sustained performance comparable to a disk with no shingled-recording. Jiguang Wan 0001, Peng Chen 0014, Changsheng Xie 0001 |
CLUSTER | 1 |
| 2012 | pCloud: An Adaptive I/O Resource Allocation Algorithm with Revenue Consideration over Public Clouds
Jianzong Wang, Daniel Gmach, Jiguang Wan 0001 |
GPC | 5 |
| 2012 | Enhancing shared RAID performance through online profilingabstractEnterprise storage systems are generally shared by multiple servers in a SAN environment. Our experiments as well as industry reports have shown that disk arrays show poor performance when multiple servers share one RAID due to resource contention as well as frequent disk head movements. We have studied IO performance characteristics of several shared storage settings of practical business operations. To avoid the IO contention, we propose a new dynamic data relocation technique on shared RAID storages, referred to as DROP, Dynamic data Relocation to Optimize Performance. DROP allocates/manages a group of cache data areas and relocates/drops the portion of hot data at a predefined sub array that is a physical partition on the top of the entire shared array. By analyzing profiling data to make each cache area owned by one server, we are able to determine optimal data relocation and partition of disks in the RAID to maximize large sequential block accesses on individual disks and at the same time maximize parallel accesses across disks in the array. As a result, DROP minimizes disk head movements in the array at run time giving rise to high IO performance. A prototype DROP has been implemented as a software module at the storage target controller. Extensive experiments have been carried out using real world IO workloads to evaluate the performance of the DROP implementation. Experimental results have shown that DROP improves shared IO performance greatly. The performance improvements in terms of average IO response time range from 20% to a factor 2.5 at no additional hardware cost. Jiguang Wan 0001, Yan Liu 0010, Qing Yang 0001, Jianzong Wang |
MSST | 1 |
| 2012 | A new high-performance, energy-efficient replication storage system with reliability guaranteeabstractIn modern replication storage systems where data carries two or more multiple copies, a primary group of disks is always up to service incoming requests while other disks are often spun down to sleep states to save energy during slack periods. However, since new writes cannot be immediately synchronized onto all disks, system reliability is degraded. This paper develops PERAID, a new high-performance, energy-efficient replication storage system, which aims to improve both performance and energy efficiency without compromising reliability. It employs a parity software RAID as a virtual write buffer disk at the front end to absorb new writes. Since extra parity redundancy supplies two or more copies, PERAID guarantees comparable reliability with that of a replication storage system. In addition, PERAID offers better write performance compared to the replication system by avoiding the classical small-write problem in traditional parity RAID: buffering many small random writes into few large writes and writing to storage in a parallel fashion. By evaluating our PERAID prototype using two benchmarks and two real-life traces, we found that PERAID significantly improves write performance and saves more energy than existing solutions such as GRAID, eRAID. Jiguang Wan 0001, Jun Wang 0001, Changsheng Xie 0001 |
MSST | 1 |
| 2012 | A Reliability Optimization Method Using Disk Reliability Degree and Data Heat DegreeabstractThe reliability of the traditional storage system can utilize data recovery and reconstruction operations to recover data in case of the disk failure; however it will result in longer data recovery time and increase the possibility of secondary failure. Hence, reliability optimization has become one of the key subjects of storage system research. In this paper we propose a novel reliability optimization method, which uses disk reliability degree to evaluate the reliability of disk based on SMART technology. And it uses data heat degree to compute the heat degree of data and the utilization degree of disk at the present time, so according to disk reliability degree and disk utilization degree we can protect current hotspots data. Data heat degree also predicts hotspots data in the future based on the current rank of data access frequency and Zipf-like distribution; we can also protect predicted new hotspots data according to disk reliability degree. Our experiment shows that disk reliability degree satisfies the actual usage and data heat degree achieves very high prediction accuracy. In order to evaluate the system performance's influence of our method, the experimental results of data migration based on RAID system demonstrate that reliability optimization method has little or no impact on the normal system performance, and outperforms the traditional reconstruct RAID system. Zhihu Tan, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 3 |
| 2010 | S2-RAID: A new RAID architecture for fast data recoveryabstractAs disk volume grows rapidly with terabyte disk becoming a norm, RAID reconstruction time in case of a failure takes prohibitively long time. This paper presents a new RAID architecture, S2-RAID, allowing the disk array to reconstruct very quickly in case of a disk failure. The idea is to form skewed sub RAIDs (S2-RAID) in the RAID structure so that reconstruction can be done in parallel dramatically speeding up data reconstruction time and hence minimizing the chance of data loss. To make such parallel reconstruction conflict-free, each sub-RAID is formed by selecting one logic partition from each disk group with size being a prime number. We have implemented a prototype S2-RAID system in Linux operating system for the purpose of evaluating its performance potential. SPC IO traces and standard benchmarks have been used to measure the performance of S2-RAID as compared to existing baseline software RAID, MD. Experimental results show that our new S2-RAID speeds up data reconstruction time by a factor of 3 to 6 compared to the traditional RAID. At the same time, S2-RAID shows similar or better production performance than baseline RAID while online RAID reconstruction is in progress. Jiguang Wan 0001, Qing Yang 0001, Changsheng Xie 0001 |
MSST | 1 |
| 2010 | Cache Blocks: An Efficient Scheme for Solid State Drives without DRAM CacheabstractMost solid state drives use DRAM for device's cache, the volatile memory provides the I/O Caching ability, and maintains the drives' mapping table (indicates the correspondence between physical unit and logical unit). However, when the drives' power shut down unexpected, the volatile DRAM memory may lose the caching data, which did not have time to write to the drives' storage media, so the dirty data generated. This paper proposes an efficient management scheme for low cost Solid State Drives, with low cost ASIC controller chip, only has internal SRAM memory, and no external DRAM. We use a kind of Cache Blocks: when write requests come, write in these areas first, and write in the sequentially physical place, and the limited internal SRAM for the mapping tables maintaining and data transferring. We propose some efficient methods: 1)using flash memory as cache, 2) page mapping for cache blocks regions, block mapping for data blocks regions, 3) binding two planes operation, 4) using the idle internal plane SRAM as data buffer to improve the I/O performance, without DRAM. So we avoid the dirty data when the power loses unexpected. And this scheme is energy-efficient and low cost. We test the scheme on our own SSD test board, with the pool SRAM size, the I/O performances do not decrease two much, and the random write even better about 20%, compare to the SSD with DRAM. The experiment also shows, this scheme cuts about 21% energy than the DRAM architecture. And it may be adapted in consumer electronics area. Fei Wu 0005, Xiang Chen 0028, Jiguang Wan 0001 |
NAS | 3 |
| 2010 | A Low Cost and Inner-round Pipelined Design of ECB-AES-256 Crypto Engine for Solid State DiskabstractSolid-State Disks (SSD) are widely used in government and security departments owing to its faster speed of data access, more durability, more shock and drop, no noise, lower power consumption, lighter weight compared with Magnetic disk. As a result, the demand of security for storing data has been generated. The Advanced Encryption Standard (AES) is today's key data encryption standard for protecting data, but the implementation of high-speed AES encryption engine needs to consume a large number of hardware resources. This paper presents a low-cost and inner-round pipelined ECB-256-AES encryption engine. Through sharing the resources between the AES encryption module and the AES decryption module and using the look-up table for the SubBytes and InvSubBytes operations, the logic resources have been largely reduced; by using loop rolling and inner-round pipelined techniques, a high throughput of encryption and decryption operations is achieved. A 1.986Gbits/s throughput and 232.748MHz clock frequency are achieved using 614 slices of the Xilinx xc6slx45-3fgg484. The simulation results show that the AES crypto design is able to meet the read and write speed of SATA 1.0 interface. Fei Wu 0005, Liang Wang 0057, Jiguang Wan 0001 |
NAS | 3 |