EDBT 2026 Demo / reviewers in the wild / expert
Si Wu 0003
dblp:25/437-0003
· DBLP profile ↗
34ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-9757-0629ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 5 first-author · 13 since 2021Computer networks · 8 · 4 first-author · 6 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedRFF: Enhanced Federated Random Fourier Feature Framework for IoT Anomaly Detection
Chaoqun Li 0002, Keyuan Qiu, Jinyao Liu, Xianglong Zhang, Huanle Zhang, Si Wu 0003, Feng Li 0002, Pengfei Hu 0001 |
ICDCS | 7 |
| 2026 | ROE: Repair-Oriented Encoding for Erasure Codes with Localities
Hongjing Yu, Si Wu 0003, Jinyao Liu, Feng Li 0002 |
INFOCOM | 2 |
| 2026 | CEAT: Context-Emotion Adversarial Training Framework for Robust Emotion-Driven Fraud DetectionabstractThe rapid proliferation of emotion-aware web services has necessitated the analysis of multimodal user interactions. However, this introduces new vulnerabilities where adversaries exploit emotional signals to circumvent fraud detection systems. Despite its improved utility, the robustness of multimodal fraud detection against emotion-driven adversarial manipulation remains significantly underexplored. Existing paradigms often treat emotional cues as static features, overlooking the adversary's capability to strategically modulate multimodal signals (e.g., facial micro-expressions, vocal intonation, and textual styles) to mimic genuine behavior. Furthermore, prevalent evaluations are typically confined to unimodal perturbations and fail to account for context-consistent, cross-modal attacks, thereby compromising system reliability in real-world deployments. To bridge this gap, we propose Context-Emotion Adversarial Training (CEAT), a robust framework designed to fortify multimodal fraud detection against emotion-based attacks. CEAT leverages a Transformer-based architecture to synergistically model emotional features (e.g., visual dynamics and acoustic prosody) alongside semantic context derived from text, yielding a unified representation. Crucially, CEAT introduces a context-aware perturbation mechanism that injects noise into the emotional latent space during training. This process preserves semantic consistency while encouraging the learning of emotion-invariant and discriminative representations. Additionally, a contrastive learning objective is integrated to maximize the distributional divergence between genuine and adversarial samples within the latent manifold. Extensive experiments on multimodal benchmarks demonstrate that CEAT significantly outperforms state-of-the-art baselines, exhibiting superior robustness under simulated emotion-driven attack scenarios. Chaoqun Li 0002, Si Wu 0003, Yuyin Ma, Jinyao Liu, Dingyi Jia, Mingda Han, Feng Li 0002, Pengfei Hu 0001 |
WWW | 2 |
| 2026 | BeeQoS: A Cloud-Native QoS System for Adaptive and Scalable Multi-Priority Bandwidth GuaranteesabstractModern cloud applications, from interative web services to mobile and WoT workloads, generate highly dynamic multi-tenant network demands. Guaranteeing priority-aware bandwidth remains challenging: legacy shapers like Linux Traffic Control Hierarchical Token Bucket are static and unscalable, while cloud-native solutions such as Cilium offer only coarse-grained rate limiting. We present BeeQoS, a cloud-native QoS system that delivers low-latency, adaptive, and scalable multi-priority bandwidth guarantees. BeeQoS consists of an eBPF-powered data plane for high-performance, fine-grained per-packet shaping, a demand-aware control plane that senses real-time flow requirements and adaptively reallocates bandwidth, and seamless Kubernetes integration for expressive policy specification and cluster-wide scalable deployment. Evaluation shows that BeeQoS scales to 1K+ flows with stable performance, boosts high-/medium-priority throughput by 14.6%/36.4%, cuts median latency by 72.4%, reduces deployment overhead, and improves video QoE by 27.3% over state-of-practice baselines. Jinyao Liu, Si Wu 0003, Chaoqun Li 0002, Hongjing Yu, Dingyi Jia, Feng Li 0002, Pengfei Hu 0001 |
WWW | 2 |
| 2026 | Optimal repair and load balance in locally repairable codes: Design and evaluation
Ximeng Chen, Si Wu 0003, Yinlong Xu 0001 |
Future Gener. Comput. Syst. | 2 |
| 2026 | Localitycache: Toward efficient straggler tolerance in LRC-coded storage via caching local parity blocksabstractModern distributed storage systems increasingly employ Locally Repairable Codes (LRCs) to provide reliable, low-cost data storage with high repair efficiency. However, the presence of stragglers, i.e., nodes that unpredictably slow down, can significantly impact access latency. Traditional approaches for handling stragglers, such as detection, blacklisting, or speculative execution, are often insufficient for efficient straggler tolerance. In this paper, we show how an in-memory caching strategy coupled with LRCs can bypass stragglers without relying on precise straggler detection. We propose LocalityCache, a novel in-memory caching mechanism designed for LRC-coded distributed storage systems, which effectively mitigates the impact of stragglers by caching local parity blocks. We provide theoretical guarantees for LocalityCache and show that caching local parity blocks minimizes the likelihood of encountering stragglers. Additionally, we devise optimized workflows for write, read, and repair operations under LocalityCache to ensure system efficiency. We implement LocalityCache in a distributed key–value store prototype atop Redis. Our extensive testbed evaluations show that LocalityCache can significantly reduce read latency of the baselines by up to 73.6% in the presence of stragglers. Ximeng Chen, Si Wu 0003, Yinlong Xu 0001 |
High Confid. Comput. | 2 |
| 2026 | UHM: Unified Transferring and Pooling Over Heterogeneous GPU MemoriesabstractWhile existing far memory and disaggregated memory solutions provide a foundation for addressing limitations of single-node memory capacity and inefficient resource allocation in data centers, they predominantly focus on host memory, overlooking the critical demands of GPU-centric workloads. A key bottleneck in scaling GPU memory is the lack of connectivity and interoperability between GPUs, which is exacerbated by their heterogeneity. To bridge this gap, this paper proposes UHM, a unified data transferring and memory pooling scheme for heterogeneous GPU memories. UHM establishes the communication channels between heterogeneous GPU/host memories and leverages double data buffers for pipelined and reliable transfer. Furthermore, UHM unifies both local and remote memories to build a memory pool. The pooling scheme effectively integrates local and remote resources, performs efficient caching management in local memory, and optimizes memory block management for remote memory resources. Evaluation on a heterogeneous GPU cluster demonstrates that UHM significantly reduces the data transfer latency (up to 87.2%), improves the cache hit ratio, reduces runtime memory allocation latency (up to 94.7%), while enhancing the overall memory utilization (24.7%). Jinyao Liu, Si Wu 0003, Chaoqun Li 0002, Shaowei Li, Hongjing Yu, Fengxi Zhou, Feng Li 0002, Xiuzhen Cheng, Pengfei Hu 0001 |
IEEE Trans. Computers | 2 |
| 2025 | TSAJS: Efficient Multi-Server Joint Task Scheduling Scheme for Mobile Edge ComputingabstractMobile Edge Computing (MEC) utilizes edge servers to offload the computational burden from cloud infrastructure. By providing low-latency and high-bandwidth services, MEC enables mobile users and IoT devices to efficiently offload and execute computational tasks at the network edge. However, optimizing communication and computational resources in a multi-user, multi-server MEC environment remains a significant challenge. In this paper, we propose TSAJS, an efficient multi-server joint task scheduling scheme designed to enhance the effectiveness of MEC offloading. We model the task offloading and resource allocation problem as a Mixed-Integer Nonlinear Programming (MINLP) problem, aiming to maximize user offloading gain by minimizing task completion time and energy consumption. A heuristic algorithm for offloading is introduced by combining threshold-triggering and simulated annealing to effectively avoid local optima and converge toward the global optimum. Meanwhile, the optimal solution for resource allocation is derived using the Karush-Kuhn-Tucker (KKT) conditions. Experimental results demonstrate that TSAJS delivers near-optimal performance, outperforming traditional methods in terms of user offloading effectiveness. Its efficiency enables solution finding within polynomial time, while also adapting to the preferences of users and service providers. Chaoqun Li 0002, Rongsheng Fan, Hesong Wang, Mingda Han, Si Wu 0003, Feng Li 0002, Pengfei Hu 0001 |
ICDCS | 5 |
| 2025 | Leveled Product Codes for Optimal Block Repairs in Geo-distributed Storage Systems
Si Wu 0003, Guantian Lin, Patrick P. C. Lee, Yinlong Xu 0001 |
INFOCOM | 1 |
| 2025 | ConfAgent: Towards Intelligent Network Configuration Via LLM AgentabstractAs network scale and complexity continue to increase, managing network configurations has become an increasingly challenging task. Existing configuration tools often depend on low-level, abstract intermediate representations, which require users to have substantial technical expertise. This reliance not only increases the learning curve but also heightens the risk of configuration errors. Recent advances in Large Language Models (LLMs) have demonstrated strong potential for automating tasks across various domains. However, their applications to network configuration generation remain limited due to several challenges, including hallucination, restricted context length, and insufficient adaptability to domain-specific requirements. To address these issues, we propose ConfAgent, an advanced network configuration generation system powered by a multi-model intelligent agent. ConfAgent comprises four key components: a conflict detector, an information extractor, a routing algorithm coder, and a formal synthesizer. These components collaborate to accurately interpret complex configuration intents, detect potential conflicts, and generate robust code and network configurations through intuitive natural language interactions. Extensive experiments conducted on the NetConfEval benchmark demonstrate that ConfAgent consistently outperforms existing state-of-the-art methods by margins ranging from 36 % to 100 %, particularly excelling in configuration tasks for large-scale network topologies. Shaowei Li, Zhiwen Gan, Jinyao Liu, Chengxi Gao, Fuliang Li, Si Wu 0003, Pengfei Hu 0001, Feng Li 0002 |
IWQoS | 6 |
| 2025 | Enabling High Performance and Resource Utilization in Clustered Cache via Hotness Identification, Data Copying, and Instance MergingabstractIn-memory cache systems such as Redis provide low-latency and high-performance data access for modern internet services. However, in large-scale Redis systems, the workloads show strong skewness and varied locality, which degrades system performance and incurs low CPU utilization. Though there are many approaches toward load imbalance, the two-layered architecture of Redis makes its workload skewness show special characteristics. Redis first maps data into data groups, which is calledGroup Mapping. Then the data groups are distributed to instances by Instance Mapping.Under Redis's layered architecture, it gives rise to a small number of hot-spot instances with very limited hot data groups, as well as a large number of remaining cold instances. To improve Redis's performance and CPU utilization, it entails the accurate identification of instance and data group hotness, and handling hot data groups and cold instances. We propose HPUCache+ to address the hot-spot problem via hotness identification, hot data copying, and cold instance merging. HPUCache+ accurately and dynamically detects instance and data group hotness based on multiple resources and workload characteristics at low cost. It enables access to multiple data copies by dynamically updating the cached mapping in Redis client, achieving high user access performance with Redis client compatibility, while providing highly self-definable service level agreement. It also proposes an asynchronous instance merging strategy based on disk snapshots and temporal caches, which separates the massive data movement from the critical user access path to achieve high-performance instance merging. We implement HPUCache+ into Redis. Experiments show that, compared to the native Redis design, HPUCache+ achieves up to 2.3$\times$and 3.5$\times$throughput gains, 11.3$\times$and 14.3$\times$CPU utilization gains, respectively. It also achieves up to 50% less CPU and 75% less memory consumption compared to the state-of-the-art approach Anna. Si Wu 0003, Zhipeng Li 0005, Yongkun Li 0001, Yinlong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2024 | Optimal Wide Stripe Generation in Locally Repairable Codes via Staged Stripe MergingabstractLarge-scale storage systems increasingly adopt era-sure coding for low-cost reliable storage by storing stripes of data and parity blocks. To further achieve higher storage savings with performance guarantees, enterprises and academia explore wide stripes with locally repairable codes (LRCs). However, how to efficiently generate wide LRC stripes remains a non-trivial issue. Current approaches on stripe merging provide a starting point for merging narrow LRC stripes into wide LRC stripes, but they lack flexibility in the number of stripes being merged and the stripe width of the wide LRC stripes being formed. In this paper, we propose staged stripe merging for wide stripe generation in LRCs, by allowing a flexible number of narrow LRC stripes to be gradually merged into a wide stripe in multiple stages with optimality guarantees. In particular, we design an optimal data placement scheme for a group of narrow LRC stripes to be merged, such that it provably minimizes the cross-cluster network traffic for stripe merging across each stage, while maintaining the repair efficiency of LRCs. We implement the optimal data placement scheme for two production LRC constructions in a distributed storage prototype. Evaluation shows that our optimal data placement scheme significantly reduces the stripe merging time compared with several baselines. Si Wu 0003, Guantian Lin, Patrick P. C. Lee, Cheng Li 0001, Yinlong Xu 0001 |
ICDCS | 1 |
| 2024 | Designing Non-uniform Locally Repairable Codes for Wide Stripes under Skewed File AccessesabstractLarge-scale storage systems increasingly deploy erasure coding to provide high-reliability and low-cost data storage. To achieve further storage savings while maintaining repair efficiency, enterprises and academia explore wide stripes using locally repairable codes (LRCs) and several constructions are proposed. However, we find that these wide stripe LRCs neglect the file access information within the stripe, which matters as these LRCs assign uniform local group sizes while a wide stripe comprises many files with non-uniform or skewed access frequencies. Consequently, the repair efficiency is compromised. To this end, we design a novel wide stripe LRC construction called non-uniform LRC by fully considering the skewed file accesses in a wide stripe. Our non-uniform LRC assigns small group sizes for hot files in the stripe to achieve low repair cost and large group sizes for cold files to realize low storage overhead. It performs well under general coding parameters, adapts to irregular file sizes, provides high fault tolerance, and can be flexibly deployed to various data placements in real-world clustered storage systems. We implement our non-uniform LRC in a distributed storage prototype. Evaluations show that it greatly reduces the degraded read time and improves the node repair rate compared with the state-of-the-art wide stripe LRCs. Guantian Lin, Si Wu 0003, Cheng Li 0001, Yinlong Xu 0001 |
ICPP | 2 |
| 2024 | Harmonizing Repair and Maintenance in LRC-Coded StorageabstractModern storage systems not only introduce data redundancy for fault tolerance, but also conduct regular main- tenance operations on storage nodes for system robustness. Erasure coding provides storage-efficient redundancy and has been widely deployed in production, yet it also incurs substantial bandwidth and I/O overhead due to the repair of storage failures. In particular, maintenance operations make storage nodes temporarily unavailable and lead to data unavailability, thereby incurring repair overhead for erasure-coded storage. In this paper, we study Locally Repairable Codes (LRCs), a class of practical repair-efficient erasure codes, and show that there exists an inherent performance trade-off between the repair and maintenance operations of LRCs in data center settings, such that the repair performance in regular (i.e., no-maintenance) and maintenance modes cannot be simultaneously optimized. To this end, we design a configurable data placement scheme that operates along the trade-off subject to fault-tolerance constraints. We prototype our data placement scheme atop Hadoop HDFS and show how it balances the performance trade-off of repair and maintenance operations in real network environments. Keyun Cheng, Si Wu 0003, Xiaolu Li 0002, Patrick P. C. Lee |
SRDS | 2 |
| 2024 | Elastic Reed-Solomon Codes for Efficient Redundancy Transitioning in Distributed Key-Value StoresabstractModern distributed key-value (KV) stores increasingly adopt erasure coding to reliably store data. To adapt to the changing demands on access performance and reliability requirements, distributed KV stores perform redundancy transitioning by tuning the redundancy schemes with different coding parameters. However, redundancy transitioning incurs extensive network I/Os, which impair the performance of distributed KV stores. We propose a new family of erasure codes, called Elastic Reed-Solomon (ERS) codes, whose primary goal is to mitigate network I/Os in redundancy transitioning. ERS codes eliminate data block relocation, while limiting network I/Os for parity block updates via the new co-design of encoding matrix construction and data placement. ERS codes achieve such gains in both forward and backward transitioning scenarios. We realize ERS codes in a distributed KV store prototype based on Memcached, and show via testbed experiments in both local and cloud environments that ERS codes significantly reduce the latency of redundancy transitioning compared with state-of-the-arts. Si Wu 0003, Zhirong Shen, Patrick P. C. Lee, Zhiwei Bai, Yinlong Xu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | CARE: A Cost-AwaRe Eviction Strategy for Improving Throughput in Cloud EnvironmentsabstractTo utilize various resources efficiently, cloud data centers usually deploy Latency-Critical (LC) and Best-Effort (BE) applications as containers in their physical machines by assigning higher priorities for LC jobs to use the resources. Due to the increase on the workload, it needs evicting some BE jobs to deprive more resources for LC jobs to meet the Quality of Service (QoS) requirements of LC jobs. Under dynamic workload settings, the frequent BE job eviction results in significant throughput degradation. To this end, we propose a novel Cost-AwaRe Eviction strategy, CARE, which takes into account both the recalculation cost and the remaining time cost of each BE job. CARE fully exploits the two types of different costs of BE jobs and chooses a BE job with the lowest overhead for eviction under different resource demands of LC jobs. Furthermore, CARE improves the throughput of BE jobs without affecting the performance of LC jobs. We prototype and implement CARE atop Kubernetes. Experimental results show that CARE achieves up to 1.70 × throughput gain compared to state-of-the-arts while its negative impact on the performance of LC jobs is negligible. Xiyuan Liang, Lulu Yao, Si Wu 0003, Yongkun Li 0001, Yinlong Xu 0001 |
ICPADS | 3 |
| 2023 | Toward Optimal Repair and Load Balance in Locally Repairable CodesabstractErasure coding is increasingly deployed in modern clustered storage systems to provide low-cost reliable storage. In particular, Locally Repairable Codes (LRCs) are a popular family of repair-efficient erasure codes that receive wide deployment in practice. In this paper, we analyze the storage process formulated as a data partitioning phase plus a node selection phase for LRCs in clustered storage systems. We show that the conventional flat partitioning and random partitioning incur significant cross-cluster repair traffic, while the random node selection causes storage and network imbalance. To this end, we design a new storage scheme composed of an optimal partitioning strategy and an enhanced node selection strategy for LRCs. Our partitioning strategy minimizes the cross-cluster repair traffic by dividing each group of blocks into the minimum number of clusters and further compactly placing the blocks. Our node selection strategy improves load balance by choosing less-loaded clusters and nodes to store blocks with potential higher access frequency at higher priority. We implement our storage scheme on a key-value store prototype atop Memcached. Evaluation on a LAN testbed shows that our scheme greatly improves the repair performance and load balance ratio compared to the baseline. Si Wu 0003, Haifeng Liu 0004, Zhixiang Tang, Xiaochun He, Yinlong Xu 0001 |
ICPP | 2 |
| 2023 | Towards High Performance and Efficient Memory Deduplication via Mixed PagesabstractLarge pages are widely supported in modern hardware and OSes to reduce the overhead of TLB misses. However, memory deduplication can be inefficient with large pages, leading to low memory utilization. To simultaneously enjoy the benefits of high performance by accessing memory with large pages (e.g., 2 MB pages) and high deduplication rate by managing memory with base pages (e.g., 4 KB pages), we proposeSmartMemoryDeduplciation (SmartMD), which is an adaptive and efficient memory management scheme via mixed pages. Specifically, we propose lightweight schemes to periodically monitor pages’ access frequency and repetition rate, and present an adaptive conversion scheme to selectively split or reconstruct large pages. We further optimize SmartMD by developing SmartMD$^{+}$, which dynamically adjusts the page scanning cycle by monitoring the TLB miss cost, and reconstructs the split large pages in an on-demand way so as to reduce the CPU overhead of SmartMD. We further implement a prototype system and conduct extensive experiments with various workloads under different system settings. Experiment results show that SmartMD and SmartMD$^{+}$can simultaneously achieve high access performance similar to systems using large pages, and achieve a deduplication rate similar to that applying aggressive deduplication scheme (i.e., KSM) on base pages. Lulu Yao, Yongkun Li 0001, Fan Guo 0003, Si Wu 0003, Yinlong Xu 0001, John C. S. Lui |
IEEE Trans. Computers | 4 |
| 2022 | DEPART: Replica Decoupling for Distributed Key-Value Storage
Yongkun Li 0001, Patrick P. C. Lee, Yinlong Xu 0001, Si Wu 0003 |
FAST | 5 |
| 2022 | Repair-Optimal Data Placement for Locally Repairable Codes with Optimal Minimum Hamming DistanceabstractModern clustered storage systems increasingly adopt erasure coding to realize reliable data storage at low storage redundancy. Locally Repairable Codes (LRC) are a family of practical erasure codes with high repair efficiency. Among various LRC constructions, Optimal-LRC is a recently proposed LRC approach that achieves the optimal Minimum Hamming Distance with low theoretical repair costs. In this paper, we consider the repair performance of Optimal-LRC in clustered storage systems. We show that the conventional flat data placement and random data placement incur substantial cross-cluster repair traffic, which impairs the repair performance. To this end, we design an optimal data placement scheme that provably minimizes the cross-cluster repair traffic, by carefully placing each group of blocks in Optimal-LRC into a minimum number of clusters subject to single-cluster fault tolerance. We implement our optimal data placement scheme on a key-value store prototype atop Memcached, and show via LAN testbed experiments that the optimal data placement significantly improves the repair performance compared to the conventional data placements. Si Wu 0003, Cheng Li 0001, Yinlong Xu 0001 |
ICPP | 2 |
| 2022 | Optimal Data Placement for Stripe Merging in Locally Repairable CodesabstractErasure coding is a storage-efficient redundancy scheme for modern clustered storage systems by storing stripes of data and parity blocks across the nodes of multiple clusters; in particular, locally repairable codes (LRC) continue to be one popular family of practical erasure codes that achieve high repair efficiency. To efficiently adapt to the dynamic requirements of access efficiency and reliability, storage systems often perform redundancy transitioning by tuning erasure coding parameters. In this paper, we apply a stripe merging approach for redundancy transitioning of LRC in clustered storage systems, by merging multiple LRC stripes to form a large LRC stripe with low storage redundancy. We show that the random placement of multiple LRC stripes that are being merged can lead to high cross-cluster transitioning bandwidth. To this end, we design an optimal data placement scheme that provably minimizes the cross-cluster traffic for stripe merging, by judiciously placing the blocks to be merged in the same cluster while maintaining the repair efficiency of LRC. We prototype and implement our optimal data placement scheme on a local cluster. Our evaluation shows that it significantly reduces the transitioning time by up to 43.2% compared to the baseline. Si Wu 0003, Qingpeng Du, Patrick P. C. Lee, Yongkun Li 0001, Yinlong Xu 0001 |
INFOCOM | 1 |
| 2022 | HPUCache: Toward High Performance and Resource Utilization in Clustered Cache via Data Copying and Instance MergingabstractAs one of the most popular in-memory cache systems, Redis provides low-latency and high-performance data access for modern internet services. However, in large-scale Redis systems, the access skewness and locality in storage workloads induce a small number of hot-spot instances with degraded system performance and massive cold instances with low CPU utilization. This paper proposes HPUCache to address the hot-spots via data copying and cold instances via instance merging. HPUCache fully utilizes the cached mapping in Redis client, and dynamically updates this mapping to enable access to the multiple data copies. Hence it can manage multiple copies achieving both Redis client compatibility and high user access performance. It also proposes an asynchronous instance merging strategy based on disk snapshots and temporal caches, which separates the massive data movement from the critical user access path to achieve high performance instance merging. We integrate HPUCache into Redis. Experiments under two types of workloads show that, compared to the native Redis design, HPUCache achieves 2.4× and 3.5× performance gains, 2× and 6× CPU utilization gains respectively. Si Wu 0003, Yongkun Li 0001, Yinlong Xu 0001, Fei Chen 0009 |
IWQoS | 2 |
| 2022 | Optimal Repair-Scaling Trade-off in Locally Repairable Codes: Analysis and EvaluationabstractHow to improve the repair performance of erasure-coded storage is a critical issue for maintaining high reliability of modern large-scale storage systems. Locally repairable codes (LRC) are one popular family of repair-efficient erasure codes that mitigate the repair bandwidth and are deployed in practice. To adapt to the changing demands of access efficiency and fault tolerance, modern storage systems also conduct frequent scaling operations on erasure-coded data. In this article, we analyze the optimal trade-off between the repair and scaling performance of LRC in clustered storage systems. Specifically, we focus on two optimal repair-scaling trade-offs, and design placement strategies that operate along the two optimal repair-scaling trade-off curves subject to the fault tolerance constraints. We prototype and evaluate our placement strategies on a LAN testbed, and show that they outperform the conventional placement schemes in repair and scaling operations. Si Wu 0003, Zhirong Shen, Patrick P. C. Lee, Yinlong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | On the Optimal Repair-Scaling Trade-off in Locally Repairable CodesabstractHow to improve the repair performance of erasure-coded storage is a critical issue for maintaining high reliability of modern large-scale storage systems. Locally repairable codes (LRC) are one popular family of repair-efficient erasure codes that mitigate the repair bandwidth and are deployed in practice. To adapt to the changing demands of access efficiency and fault tolerance, modern storage systems also conduct frequent scaling operations on erasure-coded data. In this paper, we analyze the optimal trade-off between the repair and scaling performance of LRC in clustered storage systems. Specifically, we design placement strategies that operate along the optimal repair-scaling trade-off curve subject to the fault tolerance constraints. We prototype and evaluate our placement strategies on a LAN testbed, and show that they outperform the conventional placement scheme in repair and scaling operations. Si Wu 0003, Zhirong Shen, Patrick P. C. Lee |
INFOCOM | 1 |
| 2020 | Enabling I/O-Efficient Redundancy Transitioning in Erasure-Coded KV Stores via Elastic Reed-Solomon CodesabstractModern key-value (KV) stores increasingly adopt erasure coding to reliably store data. To adapt to the changing demands on access performance and reliability requirements, KV stores perform redundancy transitioning by tuning the redundancy schemes with different coding parameters. However, redundancy transitioning incurs extensive I/Os, which impair the performance of KV stores. We propose a new family of erasure codes, called Elastic Reed-Solomon (ERS) codes, whose primary goal is to mitigate I/Os in redundancy transitioning. ERS codes eliminate data block relocation, while limiting I/Os for parity block updates via the new co-design of encoding matrix construction and data placement. We realize ERS codes as a KV store atop Memcached, and show via LAN testbed experiments that ERS codes significantly reduce the latency of redundancy transitioning compared to state-of-the-arts. Si Wu 0003, Zhirong Shen, Patrick P. C. Lee |
SRDS | 1 |
| 2020 | Popularity-Based Online Scaling for RAID Systems Under General SettingsabstractThe ever-increasing demand of storage capacity and system performance leads to the scaling requirement in redundant arrays of independent disks (RAID)-structured storage systems. Existing approaches mainly focus on minimizing data migration in offline scenario, but not consider the user accesses issued by applications. However, in online scenario, the scaling I/Os and user I/Os usually interfere with each other, and result in significant performance degradation. Thus, it is of big significance to develop an efficient online scaling scheme to mitigate the impact of I/O interference. In this paper, we propose popularity-based online scaling (POS), which exploits workload locality by scaling storage zones with high popularity in a higher priority. POS can efficiently alleviate the interference between scaling I/Os and user I/Os, and it is also general enough to scale various RAID systems like RAID-0, RAID-5, and RAID-6. It can also be readily deployed atop various conventional RAID scaling approaches to improve their performance. To demonstrate the effectiveness of POS, we implement various conventional RAID scaling approaches, and also implement POS on top of these approaches for comparison. Extensive simulations with real-world workloads show that POS can efficiently reduce the response time of user requests and scaling I/Os, and also improves the sequentiality of data accesses. Chengjin Tian, Yongkun Li 0001, Si Wu 0003, Jinzhong Chen, Liu Yuan 0002, Yinlong Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | A Hierarchical RAID Architecture Towards Fast Recovery and High ReliabilityabstractDisk failures are very common in modern storage systems due to the large number of inexpensive disks. As a result, it takes a long time to recover a failed disk due to its large capacity and limited I/O. To speed up the recovery process and maintain a high system reliability, we propose a hierarchical code architecture with erasure codes, OI-RAID, which consists of two layers of codes, outer layer code and inner layer code. Specifically, the outer layer code is deployed with disk grouping technique based on Balanced Incomplete Block Design (BIBD) or complete graph with skewed data layout to provide efficient parallel I/O of all disks for fast failure recovery, and the inner layer code is deployed within each group of disks to provide high reliability. As an example, we deploy RAID5 in both layers to achieve fault tolerance of at least three disk failures, which meets the requirement of data availability in practical systems, as well as much higher speed up ratio for disk failure recovery than existing approaches. Besides, OI-RAID also keeps the optimal data update complexity and incurs low storage overhead in practice. Yongkun Li 0001, Chengjin Tian, Si Wu 0003, Yueming Zhang, Yinlong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | DSC: Dynamic stripe construction for asynchronous encoding in clustered file systemabstractNowadays many clustered file systems adopt asynchronous encoding which transforms replicated data into erasure coding to maintain data availability with bounded storage overhead. Existing implementations of asynchronous encoding construct coding stripes with logically sequential data blocks, which suffers from heavy cross-rack traffic and necessitates data block redistribution. Recent work [12] solves this problem by carefully distributing replicated data blocks among racks at the time when they are being written, but it is not applicable to the cases when existing systems have different data layouts or the data layout changes. In this paper, we propose Dynamic Stripe Construction (DSC) to transform N-way replication to erasure coding. DSC does not induce to any cross-rack traffic for encoding, and it does not require data block redistribution after encoding. Besides, DSC is general enough to be applied to any existing CFSes with various erasure codes, and it can also be deployed on a distributed file system in a hot-plugging-in manner. To validate the effectiveness of DSC, we implement it on HDFS. Through extensive testbed experiments in a real storage cluster, we show that DSC can significantly increase the encoding throughput and reduce the foreground user response time over the traditional approach. Shuzhan Wei, Yongkun Li 0001, Yinlong Xu 0001, Si Wu 0003 |
INFOCOM | 4 |
| 2016 | OI-RAID: A Two-Layer RAID Architecture towards Fast Recovery and High ReliabilityabstractA lot of inexpensive disks in modern storage systems induce frequent disk failures. It takes a long time to recover a failed disk due to its large capacity and limited I/O. This paper proposes a hierarchical architecture of erasure code, OI-RAID. OI-RAID consists of two layers of codes, outer layer code and inner layer code. The outer layer code is based on disk grouping and Balanced Incomplete Block Design (BIBD) with skewed data layout to provide efficient parallel I/O of all disks for failure recovery. Inner layer code is deployed within a group of disks. As an example, we deploy RAID5 in both layers and present detailed performance analysis. With RAID5 in both layers, OI-RAID tolerates at least three disk failures meeting practical data availability, and achieves much higher speed up of disk failure recovery than existing approaches, while keeping optimal data update complexity and practically low storage overhead. Yinlong Xu 0001, Yongkun Li 0001, Si Wu 0003 |
DSN | 4 |
| 2016 | Improving Read Performance of SSDs via Balanced Redirected ReadabstractModern SSDs have been a competitive alternative to traditional hard disks because of higher random access performance, lower power consumption and less noise. Although SSDs usually have multiple channels with each channel connected to multiple chips to improve their performance with parallel channels and chips, some recent studies show that contentions among I/O requests make SSDs read performance degrade notoriously, even worse than random write performance. Meanwhile, MLC/TLC flash memory technology increases modern SSDs' capacity while sacrificing their reliability. So fault tolerance, such as chip-level RAID, is also a necessity for SSDs. We propose a Balanced Redirected Read (BRR), which redirects some read requests from busy chips to relatively idle chips by decoding target data chunks with the data/parity chunks in a parity group. We also design a new layout of RAID-5 in SSDs, which distributes adjacent chips to different parity groups such that the degrees of loads on different chips in a group are more likely different. So BRR has more chances to redirect read requests from busy chips to idle chips. Compared with the recently proposed Parallel Issue Queuing (PIQ) using reordering I/O requests technique, BRR schedules the requests among chips more balanced. We implement BRR atop a trace driven simulator, and extensive experiments with real-world workloads show that BRR reduces the waiting time of read requests over PIQ (FIFO) by as much as 38.4% (77.1%) and averagely 14.2% (23.8%), and BRR can also slightly improve write performance due to the improvement of read performance. Yinlong Xu 0001, Dongdong Sun, Si Wu 0003 |
NAS | 4 |
| 2016 | Efficient Parity Update for Scaling RAID-like Storage SystemsabstractIt is inevitable to scale RAID systems with the increasing demand of storage capacity and I/O throughput. When scaling RAID systems, we will always need to update parity to maintain the reliability of the storage systems. There are two schemes, read-modify-write (RMW) and read-reconstruct-write (RCW), to update parity. However most existing scaling approaches simply use RMW to update parity. While in many scenarios for existing scaling approaches, RCW performs better in terms of the number of scaling I/Os. In this paper, we propose an algorithm, called EPU, to analyze which of RCW and RMW is better for a scaling scenario and select the more efficient one to save the scaling I/Os. We apply EPU to online scaling scenarios and further use two optimizations, I/O overlap and access aggregation, to enhance the online scaling performance. Using Scale-CRS, one of the existing scaling approaches, as an example, we show via numerical studies that Scale-CRS+EPU reduces the amount of scaling I/Os over the traditional Scale-CRS in many scaling cases. To justify the online efficiency of EPU, we implement both Scale-CRS+EPU and Scale-CRS in a simulator with Disksim as a working module. Through extensive experiments, we show that Scale-CRS+EPU reduces the scaling time and the online response time of user requests over Scale-CRS. Dongdong Sun, Yinlong Xu 0001, Yongkun Li 0001, Si Wu 0003, Chengjin Tian |
NAS | 4 |
| 2016 | I/O-Efficient Scaling Schemes for Distributed Storage Systems with CRS CodesabstractSystem scaling becomes essential and indispensable for distributed storage systems due to the explosive growth of data volume. Considering that fault-protection is a necessity in large-scale distributed storage systems, and Cauchy Reed-Solomon (CRS) codes are widely deployed to tolerate multiple simultaneous node failures, this paper studies the scaling problem of distributed storage systems with CRS codes. In particular, we formulate the scaling problem with an optimization model in which both the post-scaling encoding matrix and the data migration policy are assumed to be unknown in advance. To minimize the I/O overhead, we propose a three-phase optimization scaling scheme for CRS codes. Specifically, we first derive the optimal post-scaling encoding matrix under a given data migration policy, then optimize the data migration process using the selected post-scaling encoding matrix, and finally exploit the Maximum Distance Separable (MDS) property to further optimize the designed data migration process. Our scaling scheme requires minimal data movement while achieving uniform data distribution. Moreover, it requires to read fewer data blocks than conventional minimum data migration schemes, but still guarantees the minimum amount of migrated data. To validate the efficiency of our scheme, we implement it atop a networked file system. Extensive experiments show that our scaling scheme uses less scaling time than the basic scheme. Si Wu 0003, Yinlong Xu 0001, Yongkun Li 0001, Zhijia Yang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | POS: A Popularity-based Online Scaling scheme for RAID-structured storage systemsabstractThe ever-increasing demand of storage capability leads to scaling requirement in RAID-structured storage systems. Previous approaches to RAID scaling mainly focus on minimizing data migration, without considering the user-level application accesses. However, the mixed scaling I/Os and user accesses in practical systems will interfere with each other, which results in significant performance degradation of both the data migration time and the user response time. In this paper, we divide the whole storage space into multiple zones and measure the popularity (mainly using the metric of access frequency) of each zone. Based on the measured popularity, we propose an online scheme, namely Popularity-based Online Scaling (POS), to scale RAID-structured storage systems. The main idea of POS is to scale storage areas with high popularity first so as to better exploit workload locality. POS can efficiently alleviate the performance degradation of user response time and data migration time during the scaling process. It can be readily deployed atop various conventional RAID scaling approaches to improve their performance. To evaluate the performance of POS, we implement FastScale and FastScale with POS (POS-FS) in the same system. Through extensive benchmark studies on real-system workloads, we show that POS can efficiently reduce the response time to user requests and scaling I/Os and improve the sequentiality of data accesses. Si Wu 0003, Yinlong Xu 0001, Yongkun Li 0001, Yunfeng Zhu |
ICCD | 1 |
| 2014 | Enhancing scalability in distributed storage systems with Cauchy Reed-Solomon codesabstractSystem scaling becomes essential and indispensable for distributed storage systems due to the explosive growth of data volume. As fault-protection is also a necessity in large-scale distributed storage systems, and Cauchy Reed-Solomon (CRS) codes are widely deployed to tolerate multiple simultaneous node failures, this paper studies the scaling of distributed storage systems with CRS codes. In particular, we formulate the scaling problem with an optimization model in which both the post-scaling encoding matrix and the data migration policy are assumed to be unknown in advance. To minimize the I/O overhead for CRS scaling, we first derive the optimal post-scaling encoding matrix under a given data migration policy, and then optimize the data migration process using the selected postscaling encoding matrix. Our scaling scheme requires the minimal data movement while achieving uniform data distribution. To validate the efficiency of our scheme, we implement it atop a networked file system. Extensive experiments show that our scaling scheme reduces 7.94% to 58.87%, and 39.52% on average, of the scaling time over the basic scheme. Si Wu 0003, Yinlong Xu 0001, Yongkun Li 0001, Yunfeng Zhu |
ICPADS | 1 |