VLDB 2026 Research / reviewers in the wild / expert
Jie Yao 0001
dblp:33/3197-1
· DBLP profile ↗
44ranked-venue papers
1as first author
22since 2021 · last 2026
0009-0007-6470-7063ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rearchitecting Buffered I/O in the Era of High-Bandwidth SSDs
Yekang Zhan, Tianze Wang, Zheng Peng 0017, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
FAST | 9 |
| 2026 | FlowKV: A Parallel and IO-Friendly Key-Value Store Exploiting High-Bandwidth Solid State Drives
Jiuyue Yan, Shuhe Xia, Jie Yao 0001, Qiang Cao 0001 |
IPDPS | 6 |
| 2026 | eLDPC: An Elastic and Scalable LDPC-Decoder With Early Termination by Effectively Leveraging High-Level SynthesisabstractEmerging communication and storage embrace Low-Density Parity-Check (LDPC) codes to fully exploit their physical channels. FPGA (Field-Programmable Gate Array) is widely employed to fast prototype and accelerate the LDPC decoding with high complexity. For varying channel conditions, the FGPA decoder is desired to elastically stop iteration when meeting success condition, avoiding conservatively performing a predefined and large number of iterations. However, the dynamical-execution algorithms with adjustable parameters generally are challenging for scalable decoder structure preferred to deterministic execution logic. To overcome the problem, this paper presents an elastic and scalable HLS-based FPGA LDPC decoder architecture with early-termination to achieve high throughput and flexibility. To this end, eLDPC first provides a universal operation, fully leveraging the features of HLS to efficiently implement optimized small-scale hardware units for low-level data-update operations. Second, eLDPC presents a decoding-iteration pipeline that adds a termination-check stage to terminate the following iteration for current codeword decoding. eLDPC also presents an HLS-enhanced approach to address memory access conflicts associated with the DU pipeline. Further, eLDPC extends the number of DU decoding-iteration pipelines within a single stream to decode multiple codewords in parallel. Third, eLDPC designs elastic and independent multiple decoding streams by using FIFO queues to decouple Input, Output, and a decoding unit (DU) with variable iterations while avoiding the potential blockage of the queueing. We implement and evaluate eLDPC on a Xilinx U55C. Experiments show that eLDPC outperforms recent decoders by up to 5× with the same parameter and achieves the actual decoding throughput of up to 49.5 Gbps with high scalability and flexibility. Qiang Cao 0001, Yifan Zhang 0012, Yekang Zhan, Jie Yao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Rethinking the Request-to-IO Transformation Process of File Systems for Full Utilization of High-Bandwidth SSDs
Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
FAST | 7 |
| 2025 | Repo: Proactive Swapping Exploiting Loop Patterns in Modern ApplicationsabstractModern data-intensive applications such as large language models already outrun affordable DRAM. Page swapping to fast SSDs or network-attached memory adds capacity, but existing operating system policies often struggle when an application's working set shifts, causing costly page faults and degrading performance. Meanwhile, many applications iterate over large data objects in regular loops, which is favorable for optimization. But existing eviction and prefetch policies largely miss this opportunity, because they rely on short-term recency and reactive prefetching, leading to memory thrashing and massive uncovered faults. This paper proposes Repo, a novel swap policy that identifies and exploits intrinsic loops in modern applications. Repo utilizes PEBS-based sampling and clustering for rapid and accurate delineation of loop elements, confirms loops reliably using stable load/store counts, and crucially, coordinates eviction and prefetching proactively. Experiments on real applications show that Repo reduces page faults by up to 98 % and execution time by as much as 78 % relative to state-of-the-art baselines. Qiang Cao 0001, Yekang Zhan, Jie Yao 0001 |
ICCD | 5 |
| 2025 | HeteroGNN: A Heterogeneous Stage Division Based GNN Training Framework to Maximize CPU-GPU ParallelismabstractGraph Neural Networks (GNNs) have become inevitable tools for extracting knowledge from massive topological structure data. However, experimental observation shows that existing GNN training frameworks exhibit low efficiency when performing memory-access-intensive data preparation stage upon CPU and computation-intensive model training stage upon GPU. This is largely due to data dependency restriction between the two stages and sequential execution of computation in training iterations. Based on these observations, this paper proposes HeteroGNN, an efficient GNN training framework, to maximize parallelism of GNN training upon heterogeneous CPU-GPU architecture. Specifically, HeteroGNN first proposes a data-dependency-aware stage division policy, which divides the two stages to six phases to offer inter-stage parallelism. Further, HeteroGNN establishes a fine-grained computation partition schema, which actively partitions typical GNN computing operations into multiple schedulable tasks suitable for CPU and GPU. Finally, HeteroGNN designs an adaptive task scheduler, which adaptively schedules the tasks upon six phases to maximize GPU efficiency. Experiment results demonstrate that HeteroGNN speeds up end-to-end training time up to 1.29× comparing to state-of-the-art GNN framework, PiPAD, and 1.30× to 2.06× comparing to DGL and PyG. Xiangrui Yang 0001, Yekang Zhan, Qiang Cao 0001, Jie Yao 0001 |
ICME | 6 |
| 2025 | AIS: An Active Idleness I/O Scheduler to Reduce Buffer-Exhausted Degradation of Solid-State DrivesabstractModern solid-state drives (SSDs) continue to boost storage density and I/O bandwidth at the cost of flash-access I/O latency, especially for write, hence they prevalently deploy a build-in buffer to absorb incoming writes. However, when the buffer is used up, the applications suffer from a sudden and long performance decline, i.e., buffer-exhausted degradation (BED). To holistically understand BED and recovery, we design an automated testing toolset (SSDTest) to measure six commodity NVMe SSDs and find: (1) the occurrence of the BED strictly relies on the written-data amount, (2) BED dramatically increases I/O latency of SSDs, especially write and read-after-write, (3) BED can be conditionally reduced and recovered only after a period of idle time, and (4) a read without preceding writes is largely immune to BED, but prolongs the required idle time to recover the available buffer. Furthermore, we build a black-box SSD buffer-recovery model to quantitatively characterize the idleness-recovery behaviors and design an SSD BED predictor to make BED occurrence and buffer recovery predictable. Leveraging this model, we further design an Active Idleness I/O Scheduler (AIS) with small-sized auxiliary storage to actively regulate the I/O idle-intervals to maximize the internal buffer recovery of SSD. AIS adaptively steers incoming data to the auxiliary storage to (1) strategically keep SSD idle to reduce the occurrence of BED and (2) mitigate the tail latency of SSDs caused by read-after-writes during BED. We perform extensive evaluations under a variety of workloads. The results show that AIS improves average, 99th, 99.9th, and 99.99th-percentile latencies of SSDs by up to 29.3%, 37.3%, 78.7%, and 67.2% respectively, with up to 512MB auxiliary storage. Yekang Zhan, Xiangrui Yang 0001, Haichuan Hu, Qiang Cao 0001, Yifan Zhang 0012, Jie Yao 0001 |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | RomeFS: A CXL-SSD Aware File System Exploiting Synergy of Memory-Block Dual PathsabstractCompute eXpress Link (CXL) based Solid-State Drives (CXL-SSDs), such as the Samsung CMM-H model, promise to offer CXL.mem memory and CXL.io block dual-mode interfaces. Nonetheless, whether and how cloud applications with diverse and varying access patterns benefit from such dual-mode CXL-SSD remains an open question for academia and industry. Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
SoCC | 7 |
| 2024 | HEncode: A Highly Modularized and Efficient FPGA QC-LDPC Encoder using High Level SynthesisabstractQC-LDPC (Quasi Cyclic Low-Density Parity-Check) codes, as a regular block-based code, have been preva-lently adopted in communication and storage fields to ensure high reliability and bandwidth of data channels. However, existing Field-Programmable Gate Array (FPGA) QC-LDPC encoders designed by RTL experts are generally dedicated to specialized LDPC codes and hardware platforms without flexibility and scalability. Recently, High-Level Synthesis (HLS) was introduced to compile a high-level encoding logic into Register Transfer Level (RTL) implementations, which are low performance and hardware efficiency due to the overlarge HLS-to-RTL design space, especially for large-scale FPGA hardware. This paper proposes a highly modularized and efficient FPGA QC-LDPC encoder, HEncoder, to fully leverage HLS to achieve high bandwidth, flexibility in both code parameters, and hardware efficiency. Firstly, HEncode presents an efficient Encode Block (EB) fully exploiting the FPGA LUT characteristic. Second, HEncode designs a low-level subword-encoding pipeline using multiple EBs and subword-parallel Encode Units (EU). Third, HEncoder designs an encode module with a pipelined data stream consecutively passing Input, EU array, and Output to balance bandwidths of accessing and encoding words. Finally, HEncode develops a design space analyzer to automatically determine the encoder parameters under constrained conditions to achieve high bandwidth. We implemented and evaluated HEncode on the Xilinx U50. The results show that compared to existing encoders, HEncode gains an increase of approximately 154.5× in the peak throughput and about 5.89 × in hardware efficiency to achieve the encoding throughput of 922.66 Gbps. Xiangrui Yang 0001, Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Xiaodi Tan |
ICCD | 6 |
| 2024 | SIndex: An SSD-based Large-scale Indexing with Deterministic Latency for Cloud Block StorageabstractThe Solid State Drives (SSD) based key-value stores face significant challenges in achieving deterministic access latency. We experimentally observe the long-tail latency is mainly caused by I/O blocking induced by SSD’s internal tasks. In this paper, we propose an SSD-based SIndex to store hundreds of billions of block-mapping entries for the latency-critical cloud block storage. To hide the latency fluctuations induced by garbage collection and buffer flushing, SIndex proposes an inter-SSD I/O scheduling based on read/write separation and SSD state transition, while adopting opportunistic request speculation and balancing mechanism to mitigate read disturbance and I/O contention. Moreover, SIndex introduces a write-staging buffer cache and a two-stage sync mechanism to preferentially buffer updated data before synchronizing them to SSDs. We evaluate the SIndex prototype with a variety of benchmarks and real-world traces on commodity SSDs. SIndex is demonstrated to outperform RocksDB and other approaches by up to 11.2 × in tail latency without affecting the throughput performance. Shucheng Wang, Kaiye Zhou, Zhandong Guo, Qiang Cao 0001, Jun Xu 0037, Jie Yao 0001 |
ICPP | 6 |
| 2024 | FluidKV: Seamlessly Bridging the Gap between Indexing Performance and Memory-Footprint on Ultra-Fast StorageabstractOur extensive experiments reveal that existing key-value stores (KVSs) achieve high performance at the expense of a huge memory footprint that is often impractical or unacceptable. Even with the emerging ultra-fast byte-addressable persistent memory (PM), KVSs fall far short of delivering the high performance promised by PM's superior I/O bandwidth. To find the root causes and bridge the huge performance/memory-footprint gap, we revisit the architectural features of two representative indexing mechanisms (single-stage and multi-stage) and propose a three-stage KVS called FluidKV. FluidKV effectively consolidates these indexes by fast and seamlessly running incoming key-value request stream from the write-concurrent frontend stage to the memory-efficient backend stage across an intermediate stage. FluidKV also designs important enabling techniques, such as thread-exclusive logging, PM-friendly KV-block structures, and dual-grained indexes, to fully utilize both parallel-processing and high-bandwidth capabilities of ultra-fast storage hardware while reducing the overhead. We implemented a FluidKV prototype and evaluated it under a variety of workloads. The results show that FluidKV outperforms the state-of-the-art PM-aware KVSs, including ListDB and FlatStore with different indexes, by up to 9× and 3.9× in write and read throughput respectively, while cutting up to 90% of the DRAM footprint. Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Yuxing Chen 0003, Jie Yao 0001, Anqun Pan |
Proc. VLDB Endow. | 5 |
| 2024 | Explorations and Exploitation for Parity-based RAIDs with Ultra-fast SSDsabstractFollowing a conventional design principle that pays more fast-CPU-cycles for fewer slow-I/Os, popular software storage architecture Linux Multiple-Disk (MD) for parity-based RAID (e.g., RAID5 and RAID6) assigns one or more centralized worker threads to efficiently process all user requests based on multi-stage asynchronous control and global data structures, successfully exploiting characteristics of slow devices, e.g., Hard Disk Drives (HDDs). However, we observe that, with high-performance NVMe-based Solid State Drives (SSDs), even the recently added multi-worker processing mode in MD achieves only limited performance gain because of the severe lock contentions under intensive write workloads. In this paper, we propose a novel stripe-threaded RAID architecture, StRAID, assigning a dedicated worker thread for each stripe-write (one-for-one model) to sufficiently exploit high parallelism inherent among RAID stripes, multi-core processors, and SSDs. For the notoriously performance-punishing partial-stripe writes that induce extra read and write I/Os, StRAID presents a two-stage stripe write mechanism and a two-dimensional multi-log SSD buffer. All writes first are opportunistically batched in memory, and then are written into the primary RAID for aggregated full-stripe writes or conditionally redirected to the buffer for partial-stripe writes. These buffered data are strategically reclaimed to the primary RAID. We evaluate a StRAID prototype with a variety of benchmarks and real-world traces. StRAID is demonstrated to outperform MD by up to 5.8 times in write throughput. Shucheng Wang, Qiang Cao 0001, Hong Jiang 0001, Ziyi Lu, Jie Yao 0001, Yuxing Chen 0003, Anqun Pan |
ACM Trans. Storage | 5 |
| 2023 | R-LDPC: Refining Behavior Descriptions in HLS to Implement High-throughput LDPC DecoderabstractHigh-Level Synthesis (HLS) translates high-level behavior-description to Register-Transfer Level (RTL) implemen-tation in modern Field-Programmable Gate Arrays (FPGAs), accelerating domain-specific hardware developments. Low-Density Parity-Check (LDPC), as a powerful error-correction code family, has been widely implemented in hardware for building a reliable data channel over a noisy physical channel in communication and storage applications. Leveraging HLS to fast prototype high-performance LDPC decoder is intriguing with high scalability and low hardware-dependence, but generally is sub-optimal due to the lack of accurate and precise behavior descriptions in HLS to characterize iteration- and circuit-level implementation details. This paper proposes an HLS-based QC-LDPC decoder with scalable throughput by precisely refining the LDPC behavior descriptions, R-LDPC for short. To this end, R-LDPC first adopts an HLS-based LDPC decoder microarchitecture with a module-level pipeline. Second, R-LDPC offers a multi-instance-sharing one (MSO) description to explicitly define shared parts and non-shared parts for an array of check-node updating-units (CNU), eliminating redundant function modules and addressing circuits. Third, R-LDPC designs efficient single-stage and multi-stage shifters to eliminate unnecessary bit-selection circuits. Finally, R-LDPC provides invalid-element aware loop scheduling before the compile phase to avoid some unnecessary stalls at runtime. We implement an R-LDPC decoder, compared to the original HLS-based implementation, R-LDPC reduces the hardware con-sumption up to 56%, the latency up to 67%, and the decoding throughput up to 300%. Furthermore, R-LDPC is adapted to different scales, LDPC standards, and code rates, and can achieve 9.9Gbps decoding throughput in Xilinx U50. Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Hong Jiang 0001 |
DATE | 3 |
| 2023 | UHS: An Ultra-fast Hybrid Storage Consolidating NVM and SSD in ParallelabstractNon-Volatile Memory (NVM) with persistency and near-DRAM performance has been commonly used as first-level fast storage atop Solid-State Drives (SSDs) and Hard Disk Drives (HDDs), constituting classic hierarchy architecture to achieve high cost-performance. However, such NVM/SSD tiered storage overuses primary NVM with limited actual performance and under-utilizes secondary SSD with increasing bandwidth. Besides, NVM and SSD exhibit distinguished I/O characteristics, but are complementary for different I/O patterns. This motivates us to design a superior hybrid storage to fully exploit NVM and SSD simultaneously. In this paper, we propose UHS, an Ultra-fast Hybrid Storage consolidating NVM and SSD to reap their own merits with key enabled techniques. First, UHS builds a uniform yet heterogenous block-level storage view for the upper applications, e.g., file systems or key-value stores. UHS provides static address-mapping to explicitly partition the global block-space into coarse-grain NVM-zones and SSD-zones, which mainly serve the metadata and file data respectively. Second, UHS presents a fine-grain request-level NVM buffer to dynamically absorb small file-writes in runtime and then migrates them to the SSDs in the background. Third, UHS designs I/O-affinity write allocation and hash-based buffer indexing to trade off write gain and read cost of the NVM-buffer. Finally, UHS designs a multi-thread I/O model to take full advantage of parallelism in both NVM and SSD. We implement UHS and evaluate it under a variety of workloads. The experiments show that UHS outperforms SSD, NVM, Bcache-writeback (representative hierarchy storage), and Device-Mapper (state-of-the-art hybrid storage) up to 8X, 1.5X, 3.5X, and 6X respectively. Qiang Cao 0001, Jie Yao 0001 |
DATE | 3 |
| 2023 | HF-LDPC: HLS-friendly QC-LDPC FPGA Decoder with High Throughput and FlexibilityabstractLDPC (Low-Density Parity-Check) codes have become a cornerstone of transforming a noise-filled physical channel into a reliable and high-performance data channel in communication and storage systems. FPGA (Field-Programmable Gate Array) based LDPC hardware, especially for decoding with high complexity, is essential to realizing the high-bandwidth channel prototypes. HLS (High-Level Synthesis) is introduced to speed up the FPGA development of LDPC hardware by automatically compiling high-level abstract behavioral descriptions into RTL-level implementations, but often sub-optimally due to lacking effective low-level descriptions. To overcome this problem, this paper proposes an HLS-friendly QC-LDPC FPGA decoder architecture, HF-LDPC, that employs HLS not only to precisely characterize high-level behaviors but also to effectively optimize low-level RTL implementation, thus achieving both high throughput and flexibility. First, HF-LDPC designs a multi-unit framework with a balanced I/O-computing dataflow to adaptively match code parameters with FPGA configurations. Second, HF-LDPC presents a novel fine-grained task-level pipeline with interleaved updating to eliminate stalls due to data interdependence within each updating task. HF-LDPC also presents several HLS-enhanced approaches. We implement and evaluate HF-LDPC on Xilinx U50, which demonstrates that HF-LDPC outperforms existing implementations by 4× to 84× with the same parameter and linearly scales to up to 116 Gbps actual decoding throughput with high hardware efficiency. Yifan Zhang 0012, Qiang Cao 0001, Jie Yao 0001, Hong Jiang 0001 |
ICCD | 4 |
| 2023 | PMLDS: An LSM-Tree Direct Managed Storage for Key-Value Stores on Byte-Addressable DevicesabstractExisting key-value stores (KVSs) based on log-structured merge-tree (LSM-tree) have been broadly deployed in practice to leverage characteristics of conventional block storage via file system, but lack effective exploitation for emerging byte-addressed persistent memory (PM). We reveal that these KVSs running upon existing PM-aware File systems cause inefficient PM I/O behaviors, including 1) numerous page faults, 2) I/O misaligned with cacheline, and 3) bandwidth wastage of concurrent I/O threads. To make full use of PM without major modification for existing LSM-based KVSs, this paper proposes PMLDS, a direct managed storage for LSM-tree-based KVSs directly running upon PM. PMLDS acts as a unified I/O layer to handle all requests from KVS to PM. PMLDS designs an LSM-tree-aware data layout to directly map the KVS’s persistent objects to the storage slots with fixed location and size, thus simplifying and replacing the file system’s functionality with a minor modification. To improve I/O efficiency, PMLDS further presents three key techniques: 1) pre-allocating reusable data slots to avoid page faults, 2) forcing cacheline-alignment for small requests, and 3) scheduling asynchronous I/O threads to harness PM’s limited parallelism. We implement PMLDS and evaluate it with popular RocksDB under a variety of workloads. The results show that compared to representative PM-aware file systems such as Ext4-DAX, XFS-DAX, NOVA, and WineFS, PMLDS improves the write performance of RocksDB by up to 2.1 × while reducing the read latency by 20%~50%. Ziyi Lu, Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Xiangrui Yang 0001 |
ICPP | 4 |
| 2022 | Mlog: Multi-log Write Buffer upon Ultra-fast SSD RAIDabstractParity-based RAID suffering from partial-stripe write-penalty has to introduce write buffer to fast absorb and merge incoming writes, and then flush them to RAID array in batch. However, we experimentally observe that the popular buffering mechanism as Linux RAID journal and partial parity logging (PPL) becomes a bottleneck for ultra-fast SSD-based RAID, and we further uncover that the centralized log-buffer model is the prime cause. Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Jie Yao 0001 |
ICPP | 4 |
| 2022 | StRAID: Stripe-threaded Architecture for Parity-based RAIDs with Ultra-fast SSDs
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001 |
USENIX ATC | 5 |
| 2022 | Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage ServerabstractHybrid storage servers combining solid-state drives (SSDs) and hard-drive disks (HDDs) provide cost-effectiveness and μs-level responsiveness for applications. However, observations from cloud storage system Pangu manifest that HDDs are often underutilized while SSDs are overused, especially under intensive writes. It leads to fast wear-out and high tail latency to SSDs. On the other hand, our experimental study reveals that a series of sequential and continuous writes to HDDs exhibit a periodic, staircase-shaped pattern of write latency, i.e., low (e.g., 35 μs), middle (e.g., 55 μs), and high latency (e.g., 12 ms), resulting from buffered writes within HDD’s controller. It inspires us to explore and exploit the potential μs-level IO delay of HDDs to absorb excessive SSD writes without performance degradation. We first build an HDD writing model for describing the staircase behavior and design a profiling process to initialize and dynamically recalibrate the model parameters. Then, we propose a Buffer-Controlled Write approach (BCW) to proactively control buffered writes so that low- and mid-latency periods are scheduled with application data and high-latency periods are filled with padded data. Leveraging BCW, we design a mixed IO scheduler (MIOS) to adaptively steer incoming data to SSDs and HDDs. A multi-HDD scheduling is further designed to minimize HDD-write latency. We perform extensive evaluations under production workloads and benchmarks. The results show that MIOS removes up to 93% amount of data written to SSDs, reduces average and 99 th -percentile latencies of the hybrid server by 65% and 85%, respectively. Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang, Changsheng Xie 0001 |
ACM Trans. Storage | 5 |
| 2021 | VRefine: Refining Massive Surveillance Videos for Efficient Store and Fast AnalyzingabstractUbiquitous cameras continuously produce enormous surveillance videos, largely challenging the capacity of video analytics and storage system. Although such videos are encoded and compressed by codecs to effectively reduce inter-/intra-frame redundancy at pixel level, they still consume massive storage space, thus being deleted periodically to recycle storage. To reduce hardware pressure in both efficient computation and long-term storage, we propose a video refining system, VRefine, merely retaining key contents for the surveillance videos to achieve a high storage efficiency and fast video analytics. VRefine further eliminates potential inter-/intra-frame content redundancy inherent in surveillance videos from the perspective of video analysis. Specifically, VRefine gradually reduces video size in three consecutive stages: removing all B frames and part of P frames (KStore), condensing the remainder frames based on motion vectors (CStore), and extracting object-semantics into a text database (SStore) using existing object detection models. We implement and evaluate VRefine. The experimental results show that compared with the raw surveillance video, VRefine can reduce 42.3%-94.3% storage size and shorten the analyzing time by 46.5%-95.8%, with a slight and controllable reduction in prediction accuracy (3.0%). Qiang Cao 0001, Jie Yao 0001, Changsheng Xie 0001 |
CCGRID | 3 |
| 2021 | SPMFS: A Scalable Persistent Memory File System on Optane Persistent MemoryabstractThe first commercial Non-Volatile Memory (NVM) (i.e., Intel Optane DC Persistent Memory) exhibits limited parallelism, especially for write operations, which is generally neglected by existing NVM-aware file systems. Besides, the concurrent control of file systems also limits their scalability on high-performance NVMs under mainstream multi-core architectures. To effectively exploit full parallelism inherent in both NVMs and multi-core processors to enhance the overall performance, this paper proposes a novel scalable persistent memory file system, called SPMFS. SPMFS first partitions global metadata structures of the file system into per-core structures to distribute load and relieve contention. Second, SPMFS presents a fine-grained range lock to support concurrent accesses upon a file. Finally, SPMFS designs a dedicated I/O thread pool to offer optimal parallelism inherent in underlying NVM regardless of varying user-threads. We implement an SPMFS prototype and evaluate it under a variety of workloads generated by IOtest, Filebench, FIO, and production traces from Alibaba Pangu. The experiments show that SPMFS provides better scalability, and achieves up to 2.37 × write throughput improvement over state-of-the-art kernel NVM-aware file systems (Ext4-DAX, NOVA, and PMFS) and user-space file systems (Strata and Libnvmmio), without sacrificing the read performance. Yang Yang 0068, Qiang Cao 0001, Jie Yao 0001, Weikang Kong |
ICPP | 3 |
| 2021 | Exploiting Buffered Updates for Fast Streaming Graph AnalysisabstractStreaming graph analysis extracts timely insights from evolving graphs, and has gained increasing popularity. In current practice of streaming graph analysis, incoming updates are simply cached in a buffer, until being applied onto existing graph structure to construct a new snapshot. Graph algorithms then work on the new snapshot to produce up-to-date analysis result. Nevertheless, we find that for widely used monotonic graph algorithms, the analysis process can be accelerated by preprocessing buffered updates. To this end, we propose GraPU, a streaming graph analytics system for monotonic graph algorithms. Before applying updates, GraPU preprocesses buffered updates in three consecutive stages: 1) Components-based Classification first identifies the effective graph data that are actually affected by current updates, by classifying the vertices involved in buffered updates according to the predetermined connected components in underlying graph; 2) In-buffer Precomputation generates the safe and profitable intermediate values that can be later merged onto underlying graph to facilitate convergence on new snapshots, by precomputing the values of vertices involved in buffered updates; 3) Hub-vertices Division eliminates the vertex-level load imbalance for analysis on new snapshots, by automatically identifying the high-degree vertices involved in updates and efficiently distributing their high-cost computation over multiple machines. After buffered updates are applied, GraPU calculates vertex values in new snapshots using the subgraph-centric model. GraPU further presents Load-factors Guided Balancing to achieve load balance at subgraph-level, by reassigning some vertices and edges among subgraphs beforehand. Our experimental result shows that, GraPU outperforms state-of-the-art KineoGraph by up to 20.43x. Feng Sheng, Qiang Cao 0001, Jie Yao 0001 |
IEEE Trans. Computers | 3 |
| 2020 | BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage Server
Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang |
FAST | 5 |
| 2020 | SeRW: Adaptively Separating Read and Write upon SSDs of Hybrid Storage Server in CloudsabstractNowadays, cloud providers embrace hybrid storage servers to reap both high IO performance of solid-state drives (SSDs) and low-cost of hard disk drives (HDDs). These hybrid storage servers generally employ SSDs as primary storage directly serving requests from front-end applications while using HDDs as the secondary storage to provide sufficient storage capacity. Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Puyuan Yang |
ICPP | 5 |
| 2020 | GraBi: Communication-Efficient and Workload-Balanced Partitioning for Bipartite GraphsabstractMachine Learning and Data Mining (MLDM) applications, such as recommendation and topic modeling, generally represent their input data in bipartite graphs with two disjoint vertex-subsets connected only by edges between them. Despite the prevalence of bipartite graphs, existing graph partitioning frameworks have rarely sufficiently exploited their unique structures, especially the highly lopsided subset sizes and extremely skewed vertex degrees. As a result of poor partitioning quality, problems, particularly of high communication cost and severe workload imbalance, arise during subsequent computation over these bipartite graphs in distributed environments such as datacenters or HPC systems, significantly hampering the performance of MLDM applications. Feng Sheng, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
ICPP | 4 |
| 2020 | A Fast Filtering Mechanism to Improve Efficiency of Large-Scale Video AnalyticsabstractSurveillance cameras are ubiquitous around us. Emerging full-feature object-detection models can analyze surveillance videos with high accuracy but consume much computation. Directly applying these models for practical scenarios with large-scale cameras is prohibitively expensive. This, however, is wasteful and unnecessary considering that user-defined anomalies occur rarely among these videos. Therefore, we propose FFS-VA, a multi-stage Fast Filtering Mechanism for Video Analytics, to make video analytics much cost-effective. FFS-VA filters out the frames without the user-defined events by two stream-specialized filters and a cheap full-function model, to reduce the number of frames reaching the full-feature model. FFS-VA presents a global feedback-queue approach to balance the processing speeds of different filters in intra-stream and inter-stream processes. FFS-VA designs a dynamic batch technique to achieve a trade-off between throughput and latency. FFS-VA can also efficiently scale to multiple GPUs. We evaluate FFS-VA against the state-of-the-art YOLOv3 under the same hardware and video workloads. The experimental results show that under a 12.88 percent target-object occurrence rate on two GPUs, FFS-VA can support up to 30 concurrent video streams (15× more than YOLOv3) in the online case, and obtain 10× speedup when offline analyzing a stream, with an accuracy loss of less than 2 percent. Qiang Cao 0001, Hong Jiang 0001, Wenhui Zhang 0005, Jingjun Li, Jie Yao 0001 |
IEEE Trans. Computers | 6 |
| 2020 | Batch-file Operations to Optimize Massive Files Accessing: Analysis, Design, and ApplicationabstractExisting local file systems, designed to support a typical single-file access mode only, can lead to poor performance when accessing a batch of files, especially small files. This single-file mode essentially serializes accesses to batched files one by one, resulting in a large number of non-sequential, random, and often dependent I/Os between file data and metadata at the storage ends. Such access mode can further worsen the efficiency and performance of applications accessing massive files, such as data migration. We first experimentally analyze the root cause of such inefficiency in batch-file accesses. Then, we propose a novel batch-file access approach, referred to as BFO for its set of optimized Batch-File Operations , by developing novel BFOr and BFOw operations for fundamental read and write processes, respectively, using a two-phase access for metadata and data jointly. The BFO offers dedicated interfaces for batch-file accesses and additional processes integrated into existing file systems without modifying their structures and procedures. In addition, based on BFOr and BFOw, we also propose the novel batch-file migration BFOm to accelerate the data migration for massive small files. We implement a BFO prototype on ext4, one of the most popular file systems. Our evaluation results show that the batch-file read and write performances of BFO are consistently higher than those of the traditional approaches regardless of access patterns, data layouts, and storage media, under synthetic and real-world file sets. BFO improves the read performance by up to 22.4× and 1.8× with HDD and SSD, respectively, and it boosts the write performance by up to 111.4× and 2.9× with HDD and SSD, respectively. BFO also demonstrates consistent performance advantages for data migration in both local and remote situations. Yang Yang 0068, Qiang Cao 0001, Jie Yao 0001, Hong Jiang 0001 |
ACM Trans. Storage | 3 |
| 2020 | Improving Overall Performance of TLC SSD by Exploiting Dissimilarity of Flash PagesabstractTLC flash has three types of pages to accommodate the three bits in each TLC physical cell exhibiting very different program latencies. This paper proposes PA-SSD to effectively improve the overall performance by exploiting the dissimilarity of TLC pages on program latency throughout the write request handling workflow. The main idea behind PA-SSD is to coordinately allocate the same type of pages for sub-requests of any given user write request, to mitigate the potential program latency imbalance among the sub-requests, and to schedule sub-requests according to their page-types. We achieve the PA-SSD design goal by answering three key research questions: (1) how to properly determine page-type for each user write request? (2) how to actually allocate a physical page for each sub-request with an assigned page-type from (1)? (3) how to effectively schedule the sub-requests in the chips queues when their page-types are judiciously allocated from (2)? To answer the first question, we propose seven page-type specifying schemes to investigate their effects under different workloads. We answer the second question by redesigning the page allocation strategy in TLC SSD to uniformly and sequentially determine physical pages for allocation following the internal programming process of TLC flash. Lastly, a page-type aware scheduling policy is presented to reorder the sub-requests within chips' queues. Our experiments show that PA-SSD can accelerate both the write and read performance. Particularly, our proposed queue-depth based page-type specifying scheme improves write performance by 2.6 times and read performance by 1.5 times over the conventional TLC SSD. Wenhui Zhang 0005, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Analysis of and Optimization for Write-dominated Hybrid Storage Nodes in CloudabstractCloud providers like the Alibaba cloud routinely and widely employ hybrid storage nodes composed of solid-state drives (SSDs) and hard disk drives (HDDs), reaping their respective benefits: performance from SSD and capacity from HDD. These hybrid storage nodes generally write incoming data to its SSDs and then flush them to their HDD counterparts, referred to as the SSD Write Back (SWB) mode, thereby ensuring low write latency. When comprehensively analyzing real production workloads from Pangu, a large-scale storage platform underlying the Alibaba cloud, we find that (1) there exist many write dominated storage nodes (WSNs); however, (2) under the SWB mode, the SSDs of these WSNs suffer from severely high write intensity and long tail latency. To address these unique observed problems of WSNs, we present SSD Write Redirect (SWR), a runtime IO scheduling mechanism for WSNs. SWR judiciously and selectively forwards some or all SSD-writes to HDDs, adapting to runtime conditions. By effectively offloading the right amount of write IOs from overburdened SSDs to underutilized HDDs in WSNs, SWR is able to adequately alleviate the aforementioned problems suffered by WSNs. This significantly improves overall system performance and SSD endurance. Our trace-driven evaluation of SWR, through replaying production workload traces collected from the Alibaba cloud in our cloud testbed, shows that SWR decreases the average and 99til-percentile latencies of SSD-writes by up to 13% and 47% respectively, notably improving system performance. Meanwhile the amount of data written to SSDs is reduced by up to 70%, significantly improving SSD lifetime. Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang |
SoCC | 6 |
| 2019 | SPA-SSD: Exploit Heterogeneity and Parallelism of 3D SLC-TLC Hybrid SSD to Improve Write PerformanceabstractTo address the write performance problem suffered by MLC/TLC flash, researchers have proposed hybrid SSD that aims to combine the strengths of SLC flash, used as the write-buffer zone for its superior write performance, and MLC/TLC flash, as the capacity zone for its high storage density. While leveraging SLC as a physical write-buffer zone is proven effective in traditional 2D hybrid SSDs, how to effectively incorporate SLC into a 3D-stacked TLC to form a hybrid SSD has not been studied to the best of our knowledge. Yet this is a timely and important performance issue for 3D-stacked TLC given its one-shot programming scheme that results in much worse write performance than the programming scheme in 2D TLC where pages are associated with different bits of a cell and programmed in sequence separately. We believe that naively adopting the two-physical-zone approach to 3D hybrid SSD will miss a great opportunity for performance optimization because it ignores the inherent four-level parallelism (channel/chip/die/plane) of the flash chip array. To this end, we propose in this paper an SLC and Parallelism Aware hybrid SSD (SPA-SSD) to take full advantages of SLC's superior write performance, the internal multi-level parallelism of SSD, and the high storage density of 3D-stacked TLC flash. Two novel techniques enable SPA-SSD to be highly effective: (1) Type-Parallelism Joint Page Allocation (TPJ-PA), which allocates pages for write transactions according to not only available SLC pages but also parallelism to maximize resource utilization within the hybrid SSD, and (2) Queue-length and Parallelism Constrained Data Migration (QPC-DM), which triggers data migration without degrading user write performance by analyzing the device queue length and available flash resources. To evaluate performance of SPA-SSD, a hybrid SSD simulator, called HybridSim, is developed based on MQSim. Experimental results on HybridSim show that TPJ-PA improves write throughput by 60%, while QPC-DM improves write throughput by up to 10 times. Besides, trace-driven experiments on HybridSSD demonstrate that SPA-SSD improves the write latency to the flash by up to two orders of magnitude over the state-of-the-art designs. Wenhui Zhang 0005, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang |
ICCD | 4 |
| 2019 | VScan: Efficiently Analyzing Surveillance Videos via Model-joint MechanismabstractIdentifying key scenes in massive surveillance videos is extremely challenging because these scenes occur rarely while automotive identification using full-feature neural network (NN) models consumes immense computational resources. This paper proposes VScan, an efficient model-joint mechanism that adaptively schedules streams on a light-weight NN model and a full-feature NN model for analyzing videos concurrently. These two combined models with overlapped detectable objects are generic and well-developed. The former model fast scans videos to seek potential interest scenes. Only the streams with identified scenes are further analyzed by the latter model. We provide a model selection approach to select a light-weight model with an appropriate accuracy and high throughput. VScan further determines key parameters to correct predictions at runtime, thus guaranteeing the recall of target scenes. The full-feature model is responsible for ensuring output precision. To maintain a high hardware efficiency and utilization dynamically, VScan uses automatic sampling to reduce unnecessary computations, proposes stream scheduling to maximize hardware usage, and designs GPU scheduling to optimize the data processing flow. Experimental results show that benefitting from the model-joint mechanism and runtime scheduling optimizations, VScan significantly boosts the video processing throughput by up to 15x without key scene loss. Qiang Cao 0001, Jie Yao 0001, Puyuan Yang |
ICPP | 3 |
| 2019 | BFO: Batch-File Operations on Massive Files for Consistent Performance ImprovementabstractExisting local file systems, designed to support a typical single-file access pattern only, can lead to poor performance when accessing a batch of files, especially small files. This single-file pattern essentially serializes accesses to batched files one by one, resulting in a large number of non-sequential, random, and often dependent I/Os between file data and metadata at the storage ends. We first experimentally analyze the root cause of such inefficiency in batch-file accesses. Then, we propose a novel batch-file access approach, referred to as BFO for its set of optimized Batch-File Operations, by developing novel BFOr and BFOw operations for fundamental read and write processes respectively, using a two-phase access for metadata and data jointly. The BFO offers dedicated interfaces for batch-file accesses and additional processes integrated into existing file systems without modifying their structures and procedures. We implement a BFO prototype on ext4, one of the most popular file systems. Our evaluation results show that the batch-file read and write performances of BFO are consistently higher than those of the traditional approaches regardless of access patterns, data layouts, and storage media, with synthetic and real-world file sets. BFO improves the read performance by up to 22.4× and 1.8× with HDD and SSD respectively; and boosts the write performance by up to 111.4× and 2.9× with HDD and SSD respectively. BFO also demonstrates consistent performance advantages when applied to four representative applications, Linux cp, Tar, GridFTP, and Hadoop. Yang Yang 0068, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang |
MSST | 5 |
| 2019 | LT-TCO: A TCO Calculation Model of Data Centers for Long-Term Data PreservationabstractData centers have been becoming public utilities to provide large-scale computing and storage services. The Total Cost of Ownership (TCO) models for such data centers are paramount to deeply understand their cost of investment and maintenance, the cost composition of internal components, and further cost optimization directions. Existing data center TCO models focus on either high-performance data centers or key subsystems such as IT facility, lacking of holistic analysis of the data centers designed for long-term data preservation. The long-term data centers can be built with different combinations of storage media such as HDDs, tapes, and optical discs. Meanwhile, during the long operation period, devices replacement and data migration are necessary and are not negligible in cost. In order to comprehensively and quantitatively understand the cost of long-term data preservations, we proposed LT-TCO, a TCO calculation model for data centers over time. LT-TCO simulates the construction and operation of a data center to calculate the expenditure of each year. It also introduces the cost of devices replacement and data migration during the long running period. Based on the storage media as optical discs, tapes, HDDs, and SSDs, LT-TCO evaluates the corresponding capital and operational expenditure under different developing rates. The simulation result shows that in long-term preservation, data migration cost takes more than 96% of the operational expenditure. And the TCO of optical disc data centers could be the least among four storage media. Wenrui Yan, Jie Yao 0001, Qiang Cao 0001, Yifan Zhang 0012 |
NAS | 2 |
| 2018 | HPDV: A Highly Parallel Deduplication Cluster for Virtual Machine ImagesabstractData deduplication has been widely introduced to effectively reduce storage requirement of virtual machine (VM) images running on VM servers in the virtualized cloud platforms. Nevertheless, the existing state-of-the-art deduplication for VM images approaches can not sufficiently exploit the potential of underlying hardware with consideration of the interference of deduplication on the foreground VM services, which could affect the quality of VM services. In this paper, we present HPDV, a highly parallel deduplication cluster for VM images, which well utilizes the parallelism to achieve high throughput with minimum interference on the foreground VM services. The main idea behind HPDV is to exploit idle CPU resource of VM servers to parallelize the compute-intensive chunking and fingerprinting, and to parallelize the I/O-intensive fingerprint indexing in the deduplication servers by dividing the globally shared fingerprint index into multiple independent sub-indexes according to the operating systems of VM images. To ensure the quality of VM services, a resource-aware scheduler is proposed to dynamically adjust the number of parallel chunking and fingerprinting threads according to the CPU utilization of VM servers. Our evaluation results demonstrate that compared to a state-of-the-art deduplication system for VM images called Light, HPDV achieves up to 67% deduplication throughput improvement. Qiang Cao 0001, Jianzhong Huang 0001, Jie Yao 0001, Changsheng Xie 0001 |
CCGrid | 4 |
| 2018 | GraPU: Accelerate Streaming Graph Analysis through Preprocessing Buffered UpdatesabstractStreaming graph analysis extracts timely insights from evolving graphs, and has gained increasing popularity. For current streaming graph analytics systems, incoming updates are simply cached in a buffer, until being applied onto existing graph structure to construct a new snapshot. Iterative graph algorithms then work on the new snapshot to produce up-to-date analysis result. Nevertheless, we find that for widely used monotonic graph algorithms, the buffered updates can be effectively preprocessed to achieve fast and accurate analysis on new snapshots. Feng Sheng, Qiang Cao 0001, Haoran Cai, Jie Yao 0001, Changsheng Xie 0001 |
SoCC | 4 |
| 2018 | FFS-VA: A Fast Filtering System for Large-scale Video AnalyticsabstractSurveillance video cameras are ubiquitous around us. Full-feature object-detection models such as YOLOv2 can automatically analyze surveillance videos in real-time with high accuracy while consuming huge computational resources. Directly applying these models for practical scenarios with large-scale deployed cameras requires prohibitively expensive computation. This, however, is both wasteful and unnecessary considering the fact that the concerned anomalous events occur rarely among these massive volumes of video streams. Therefore, in this paper, we propose a Fast Filtering System for Video Analytics (FFS-VA), a pipelined multi-stage video analyzing system, to make video analytics much cost-effective. FFS-VA is designed to filter out vast but non-target-object frames by two prepositive stream-specialized filters and a small full-function tiny-YOLO model, to drastically reduce the number of video frames arriving at the full-feature model in the back-end. FFS-VA presents a global feedback-queue mechanism to balance the processing rates of different filters in both intra-stream and inter-stream processes. FFS-VA also designs a dynamic batch technique to achieve an adjustable trade-off between throughput and latency. FFS-VA reasonably distributes all tasks on CPUs and GPUs to fully exploit the underlying hardware resources. We implement a FFS-VA prototype and evaluate FFS-VA against the state-of-the-art YOLOv2 under the same hardware and representative video workloads. The experimental results show that under a 10% target-object occurrence rate on two GPUs, FFS-VA can support up to 30 concurrent video streams (7x more than YOLOv2) in the online case, and obtain 3x speedup when offline analyzing a stream, with an accuracy loss of less than 2%. Qiang Cao 0001, Hong Jiang 0001, Wenhui Zhang 0005, Jingjun Li, Jie Yao 0001 |
ICPP | 6 |
| 2018 | PA-SSD: A Page-Type Aware TLC SSD for Improved Write/Read Performance and Storage EfficiencyabstractTLC flash has three types of pages to accommodate the three bits in each TLC physical cell exhibiting very different program latencies, LSB (fast), CSB (medium), and MSB (slow). Conventional TLC SSD designs on page allocation to write requests do not take page types and their latency difference into consideration, missing on an important opportunity to exploit the potentials of fast writes. Wenhui Zhang 0005, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
ICS | 4 |
| 2018 | GreenSprint: Effective Computational Sprinting in Green Data CentersabstractComputational Sprinting has proven to be an effective way to boost the computing performance for bursty workloads, which allows a chip to exceed its power and thermal limits temporarily by turning on all processor cores and absorbing the extra heat dissipation with certain phase-changing materials. However, extra power available for sprinting is constrained by existing power distribution infrastructures. Using batteries alone to provide the additional power to achieve performance target not only limits the effectiveness of sprinting, but also negatively impacts the lifetime of the batteries. Leveraging renewable power supply in a green data center provides an opportunity to exploit the maximal potential of Computational Sprinting. However, the intermittent nature of renewable energy makes it very challenging. In this paper, we propose GreenSprint, a renewable energy driven approach that enables a data center to boost its computing performance efficiently by conducting computational sprinting. We present four sprinting strategies to address the challenge imposed by the intermittent and time-varying nature of renewable energy supply. We build an experimental prototype to evaluate GreenSprint on a cluster of 10 servers with a simulated solar power generator. The results show that renewable energy by itself can sustain different duration lengths of sprinting when its supply is sufficient and can improve performance by up to 4.8x for representative interactive applications. We also show the effectiveness of core-count and frequency scaling in the presence of varied renewable power and limited battery energy. Haoran Cai, Qiang Cao 0001, Hong Jiang 0001, Feng Sheng, Xiandong Qi, Jie Yao 0001, Changsheng Xie 0001, Liang Xiao 0008, Liang Gu |
IPDPS | 7 |
| 2018 | ROS: A Rack-based Optical Storage System with Inline Accessibility for Long-Term Data PreservationabstractThe combination of the explosive growth in digital data and the demand to preserve much of these data in the long term has made it imperative to find a more cost-effective way than HDD arrays and a more easily accessible way than tape libraries to store massive amounts of data. While modern optical discs are capable of guaranteeing more than 50-year data preservation without media replacement, individual optical discs’ lack of the performance and capacity relative to HDDs or tapes has significantly limited their use in datacenters. This article presents a Rack-scale Optical disc library System, or ROS in short, which provides a PB-level total capacity and inline accessibility on thousands of optical discs built within a 42U Rack. A rotatable roller and robotic arm separating and fetching discs are designed to improve disc placement density and simplify the mechanical structure. A hierarchical storage system based on SSDs, hard disks, and optical discs is proposed to effectively hide the delay of mechanical operation. However, an optical library file system (OLFS) based on FUSE is proposed to schedule mechanical operation and organize data on the tiered storage with a POSIX user interface to provide an illusion of inline data accessibility. We further optimize OLFS by reducing unnecessary user/kernel context switches inheriting from legacy FUSE framework. We evaluate ROS on a few key performance metrics, including operation delays of the mechanical structure and software overhead in a prototype PB-level ROS system. The results show that ROS stacked on Samba and FUSE as network-attached storage (NAS) mode almost saturates the throughput provided by underlying samba via 10GbE network for external users, as well as in this scenario provides about 53ms file write and 15ms read latency, exhibiting its inline accessibility. Besides, ROS is able to effectively hide and virtualize internal complex operational behaviors and be easily deployable in datacenters. Wenrui Yan, Jie Yao 0001, Qiang Cao 0001, Changsheng Xie 0001, Hong Jiang 0001 |
ACM Trans. Storage | 2 |
| 2017 | ROS: A Rack-based Optical Storage System with Inline Accessibility for Long-Term Data PreservationabstractThe combination of the explosive growth in digital data and the need to preserve much of this data in the long term has made it an imperative to find a more cost-effective way than HDD arrays and more easily accessible way than tape libraries to store massive amounts of data. While modern optical discs are capable of guaranteeing more than 50-year data preservation without migration, individual optical disks' lack of the performance and capacity relative to HDDs or tapes has significantly limited their use in datacenters. This paper presents a Rack-scale Optical disc library System, or ROS in short, that provides a PB-level total capacity and inline accessibility on thousands of optical discs built within a 42U Rack. A rotatable roller and robotic arm separating and fetching the discs are designed to improve disc placement density and simplify the mechanical structure. A hierarchical storage system based on SSD, hard disks and optical discs are presented to hide the delay of mechanical operation. On the other hand, an optical library file system is proposed to schedule mechanical operation and organize data on the tiered storage with a POSIX user interface to provide an illusion of inline data accessibility. We evaluate ROS on a few key performance metrics including operation delays of the mechanical structure and software overhead in a prototype PB-level ROS system. The results show that ROS stacked on Samba and FUSE can provide almost 323MB/s read and 236MB/s write throughput, about 53ms file write and 15ms read latency via 10GbE network for external users, exhibiting its inline accessibility. Besides, ROS is able to effectively hide and virtualize internal complex operational behaviors and be easily deployable in datacenters. Wenrui Yan, Jie Yao 0001, Qiang Cao 0001, Changsheng Xie 0001, Hong Jiang 0001 |
EuroSys | 2 |
| 2017 | Laro: Lazy repartitioning for graph workloads on heterogeneous clustersabstractDistributed graph processing frameworks attempt to eliminate workload imbalance among computing nodes. However, this expectation is generally challenged by underlying heterogeneous nodes and fluctuating graph workloads at runtime. This paper proposes Laro, a graph processing system using dynamic graph repartitioning that collects the actual processing times from all nodes, then reconstructs a vertices distribution with minimal migration costs in every iteration. We manifest that the Variation Coefficient of processing times is a critical metric to quantitatively characterize the workload imbalance among nodes at each iteration. Laro also presents a lazy repartitioning algorithm to improve migration efficiency. Laro has been implemented by extending GPS, a popular repartitioning-featured graph processing system. Our evaluation using real-world graphs shows that, by achieving more balanced workload distributions at runtime, Laro derives maximal speedup of 1.82x and 1.41x over the static Skewed Hash and the dynamic GPS respectively. Feng Sheng, Qiang Cao 0001, Haoran Cai, Jie Yao 0001, Changsheng Xie 0001 |
IPCCC | 4 |
| 2016 | Montgolfier: Latency-aware power management system for heterogeneous serversabstractHeterogeneous servers have long been introduced to improve energy efficiency in warehouse-scale computers(WSCs). However, running latency-critical web-services on heterogeneous servers is still challenging because the overheads of transition between such servers heavily impact overall benefits and performance. We propose Montgolfier, a runtime power management system based on a latency-aware feedback control mechanism. It consolidates wimpy and brawny servers into composite nodes to improve energy efficiency while ensuring QoS for latency-critical applications. Montgolfier effectively mitigates the effect of transition overhead between servers with dynamically load prediction and accurately provides thin-provisioned configurations in fine-grain manner for fluctuating loads. Our evaluation results show that Montgolfier reduces energy consumption by up to 34.9% without violating any QoS constraints. Haoran Cai, Qiang Cao 0001, Feng Sheng, Manyi Zhang, Chuanyi Qi, Jie Yao 0001, Changsheng Xie 0001 |
IPCCC | 6 |
| 2016 | Elastic-RAID: A New Architecture for Improved Availability of Parity-Based RAIDs by Elastic MirroringabstractIn this paper, we propose Elastic-RAID, a new RAID architecture to achieve high performance and high reliability for large-scale distributed and parallel storage systems. The key idea behind Elastic-RAID is to smartly utilize the free space existing in parity-based disk arrays to store additional mirroring data. This additional mirroring data redundancy, when strategically and judiciously activated and exploited in a RAID system, enables improved system I/O performance, fault tolerance and recovery. Depending on the amount of free space available and whether the emphasis is on performance or reliability, the elasticity in Elastic-RAID is manifested in how each design objective is achieved. For the performance objective, Elastic-RAID improves small-write performance by writing original and mirroring data synchronously and leaving the costly parity update in the background at a later idle/lightly-loaded time. For the reliability objective, at least two concurrent disk failures can be tolerated when Elastic-RAID is employed in a RAID5 system that has 50 percent or more free space. Higher reliability is provided for important data when free space is less than 50 percent. To achieve the design goal of elasticity, we introduce a novel data layout and addressing scheme. Our extensive trace-driven evaluations on an Elastic-RAID prototype in the typical configurations of RAID5 show that Elastic-RAID boosts the small-write performance in the normal operational state by at least 40 percent, improves the user I/O performance in the reconstruction state by at least 30 percent and shortens the recovery time by at least 40 percent. Jie Yao 0001, Hong Jiang 0001, Qiang Cao 0001, Lei Tian 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | EOPC: A parallel coding algorithm for XOR-based RAID-6 codesabstractWhile inheriting from RAID-6 codes protecting data against two simultaneous disk failures, XOR-based RAID-6 codes are low computational complexity due to only using exclusive-or operations to encode and decode, and are extensively studied and employed in practical. But the potential parallelism of these codes have not yet been sufficiently explored. In this paper, we observe that for XOR-based RAID-6 coding procedures, calculations of parity check equations can be decomposed into pre-calculating and recursive resolution phases. Moreover, these pre-calculating phases of equations can execute in parallel to obtain intermediate blocks that are further used to recursively resolve all missing blocks in a specific sequence. Based on this observation, we present a parallel coding algorithm, called EOPC, for XOR-based RAID-6 codes with the z-turn property, where there exists at least one parity check equation having only one unavailable block under their fault tolerance. We further build EOPC based on two representative XOR-based RAID-6 codes-RDP code and P-Code, to evaluate the effectiveness of EOPC. Experiment results show that EOPC approach outperforms the corresponding serialized approach by more than 50% in encoding/ decoding throughput. Wenhui Zhang 0005, Qiang Cao 0001, Shishi Tan, Jie Yao 0001 |
NAS | 5 |