VLDB 2026 Research / reviewers in the wild / expert
Chentao Wu
dblp:55/8165
· DBLP profile ↗
123ranked-venue papers
9as first author
78since 2021 · last 2026
0000-0002-6882-3754ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 8 first-author · 34 since 2021Computer networks · 19 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 14 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Security and privacy · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM Serving
Yunfei Gu, Liqiang Zhang 0010, Chentao Wu, Guangtao Xue, Jie Li 0002, Minyi Guo |
FAST | 4 |
| 2026 | AOEH: An Efficient Extendable Hashing to Reduce Read/Write Amplification for Persistent Memory
Yunfei Gu, Chentao Wu, Jie Li 0002, Junzhe Lv |
ICDE | 4 |
| 2026 | Multi-Layer Scheduling in Gig Platforms Using a Generative Diffusion Model With Duality Guidance
Zhanbo Feng, Jiong Lou, Chentao Wu, Guangtao Xue, Wei Zhao 0001, Jie Li 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | FIFO-MEP: An Efficient Multi-Eviction-Point FIFO Cache with Stable Demotion for Burst-Oriented Access MitigationabstractCaching technology is widely used in multiple areas particularly in distributed computing, where its performance is highly dependent on the cache efficiency. The cache eviction algorithm serves as the core component of a cache, primarily aimed at improving cache efficiency by reducing the cache miss ratio. Numerous eviction algorithms are proposed in recent decades and state-of-the-art methods tend to adopt lazy promotion and quick demotion designs. Lazy promotion simplifies cache-hit operations for higher throughput, while quick demotion effectively filters the low-popularity objects. However, the two designs either fail to identify burst objects or suffer from unstable demotion precision. In order to address the above problems, we propose FIFO-MEP, an efficient FIFO cache with Multiple Eviction Points. The key design of FIFO-MEP is to introduce multiple fixed-position eviction points near the head of a FIFO queue. These eviction points enable repeated inspections of objects, leading to effective identification of burst objects. Meanwhile, by fixing positions of these eviction points, FIFO-MEP delivers stable demotion precision. We implement FIFO-MEP using libCacheSim and evaluated it on 5439 production traces for three typical cache sizes, and further verify its efficiency based on Memcached. The evaluation results show that FIFO-MEP reduces the miss ratio by an average of 15.8 % across all experimental configurations. Compared to the state-of-the-art S3-FIFO, FIFO-MEP achieves cache efficiency improvement by up to 21.8 % for large cache sizes. Furthermore, FIFO-MEP yields the best performance under 51 % of all tested conditions. Ranhao Jia, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo, Liqiang Zhang 0010 |
CLUSTER | 3 |
| 2025 | MemSeer: Leverage Memory Failure Distinctions and Multi-Grained Prediction in Ultra-Scale Heterogeneous X86/ARM ClustersabstractIn high-performance ultra-scale cloud computing, heterogeneous clusters consisting of x86 and ARM architecture platforms have become increasingly common to boost performance and energy efficiency. Ensuring high availability in these environments is crucial for meeting service-level agreements. However, DRAM failures, a primary cause of server downtimes, present significant challenges to reliability, availability, and serviceability. This paper provides an in-depth analysis of memory failure characteristics across cross-architecture platforms in large-scale heterogeneous clusters. We introduce MemSeer, an AIOps-integrated tool that utilizes a multi-grained memory failure prediction approach for x86/ARM heterogeneous clusters. MemSeer improves the F1-score by 17.3% and increases recall by an average of $27 \%$ across different lead times compared to state-of-the-art methods. These advancements show great promise in reducing memory failures in cluster environments, decreasing VM interruptions by up to 42.7% and averaging 24.2% in real-world implementations. Yunfei Gu, Chentao Wu, Jieru Zhao, Jie Li 0002, Minyi Guo, Wengui Zhang, Feilong Lin |
DAC | 5 |
| 2025 | CXL-ECC: an Efficient LRC-based on-CXL-Memory-eXpander-Controller ECC to Enhance Reliability and Performance of DRAM Error CorrectionabstractCompute eXpress Link (CXL) offers an effective interface for connecting CPUs with external computing and memory devices. CXL Memory eXpander Controller (CXL-MXC) is gaining attention for its ability to boost memory capacity and bandwidth more efficiently than traditional DDR DIMMs. Despite extensive research on MXC performance and adaptation, DRAM reliability in CXL architecture remains underexplored. Traditional fault tolerance mechanisms like replica or RAID-based systems would significantly increase bandwidth overhead in the CXL fabric, adversely affecting system performance. To address this, we propose the on-CXL-Memory-Expander-Controller ECC (CXL-ECC), by using Locally Recoverable Codes (LRC) as the Inter-Channel-ECC (IC-ECC) and offloading its process to the expander, we eliminate extra memory access requests in the CXL fabric. Consequently, we conduct several experiments to demonstrate that our approach enhances DRAM reliability by more than $10^{9}$, compared to state-of-the-art ECC methods. Relative to RAID-enabled CXL switch, it reduces additional bandwidth overhead from 63.5% to 3.4% and improves system performance by 12%. Yunfei Gu, Junhao Dai, Chentao Wu, Xinfei Guo, Jieru Zhao, Jie Li 0002, Minyi Guo |
DAC | 5 |
| 2025 | STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Supportabstract3D Gaussian Splatting (3DGS) has gained popularity for its efficiency and sparse Gaussian-based representation. However, 3DGS struggles to meet the real-time requirement of 90 frames per second (FPS) on resource-constrained mobile devices, achieving only 2 to 9 FPS. Existing accelerators focus on compute efficiency but overlook memory efficiency, leading to redundant DRAM traffic. We introduce STREAMINGGS, a fully streaming 3DGS algorithm-architecture co-design that achieves fine-grained pipelining and reduces DRAM traffic by transforming from a tile-centric rendering to a memory-centric rendering. Results show that our design achieves up to 45.7 × speedup and 62.9 × energy savings over mobile Ampere GPUs. Chenqi Zhang 0002, Yu Feng 0007, Jieru Zhao, Guangda Liu, Wenchao Ding 0001, Chentao Wu, Minyi Guo |
DAC | 6 |
| 2025 | Gaze into the Pattern: Characterizing Spatial Patterns with Internal Temporal Correlations for Hardware PrefetchingabstractHardware prefetching is one of the most widely-used techniques for hiding long data access latency. To address the challenges faced by hardware prefetching, architects have proposed to detect and exploit the spatial locality at the granularity of spatial region. When a new region is activated, they try to find similar previously accessed regions for footprint prediction based on system-level environmental features such as the trigger instruction or data address. However, we find that such context-based prediction cannot capture the essential characteristics of access patterns, leading to limited flexibility, practicality and suboptimal prefetching performance. In this paper, inspired by the temporal property of memory accessing, we note that the temporal correlation exhibited within the spatial footprint is a key feature of spatial patterns. To this end, we propose Gaze, a simple and efficient hardware spatial prefetcher that skillfully utilizes footprint-internal temporal correlations to efficiently characterize spatial patterns. Meanwhile, we observe a unique unresolved challenge in utilizing spatial footprints generated by spatial streaming, which exhibit extremely high access density. Therefore, we further enhance Gaze with a dedicated two-stage approach that mitigates the over-prefetching problem commonly encountered in conventional schemes. Our comprehensive and diverse set of experiments show that Gaze can effectively enhance the performance across a wider range of scenarios. Specifically, Gaze improves performance by $\mathbf{5. 7 \%}$ and 5.4% at single-core, 11.4% and $\mathbf{8. 8 \%}$ at eight-core, compared to most recent low-cost solutions PMP and vBerti. Zixiao Chen, Chentao Wu, Yunfei Gu, Ranhao Jia, Jie Li 0002, Minyi Guo |
HPCA | 2 |
| 2025 | EACC: Efficient Agent Context Cache Sharing for Multi-Agent Systems
Sihao Cheng, Yunfei Gu, Chentao Wu |
ICA3PP (3) | 3 |
| 2025 | Generative Diffusion Model-based Energy Management in Networked Energy SystemsabstractIn recent years, the proliferation of renewable energy sources has heightened the focus on networked energy systems. These systems face significant challenges due to the unpredictable nature of energy generation and consumption, as well as the complexity of managing numerous components and parameters. To address the challenges associated with the time-consuming nature of optimization problems and the expansive solution space, we propose an innovative energy management method based on a generative diffusion model applicable to general networked energy systems. This approach aims to balance energy supply and demand while minimizing transmission costs. The efficacy of this method is validated through evaluations on real-world datasets and simulations, demonstrating a 26.6% cost reduction compared to the state-of-the-art model and a 62.8% decrease in execution time compared to existing optimizers. This research highlights the potential of generative diffusion techniques in networked energy management. Code: https://github.com/gale13/GEM. Zhanbo Feng, Jiawei Sun 0001, Jiong Lou, Chentao Wu, Wugedele Bao, Jie Li 0002 |
ICASSP | 5 |
| 2025 | Variational Perturbation Personalized Federated Learning via Prior-Posterior DistanceabstractPersonalized Federated Learning (pFL) mitigates the impact of statistical heterogeneity on FL architecture to some extent by allowing participants to use personalized models based on local data distributions. The existing pFL methods optimize from the perspective of model structure, attempting to adopt strategies that maintain model processing or quickly adapt to local data distribution capabilities. Our proposed method draws inspiration from the concept of variational inference, guiding model updates by comparing prior and posterior data distributions, and innovatively applying model variational perturbations to improve robustness. Finally, we conducted multidimensional experiments and the results show that our method outperforms the current baseline. Code: https://github.com/RezinChow/VPFL. Hefeng Zhou, Jun Wang 0012, Jiong Lou, Wugedele Bao, Chentao Wu, Jie Li 0002 |
ICASSP | 6 |
| 2025 | LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsabstractThe widespread deployment of cameras has led to an exponential increase in video data, creating vast opportunities for applications such as traffic management and crime surveillance. However, querying specific objects from large-scale video datasets presents challenges, including (1) processing massive and continuously growing data volumes, (2) supporting complex query requirements, and (3) ensuring low-latency execution. Existing video analysis methods struggle with either limited adaptability to unseen object classes or suffer from high query latency. In this paper, we present LOVO, a novel system designed to efficiently handle compLex Object queries in large-scale VideO datasets. Agnostic to user queries, LOVO performs one-time feature extraction using pre-trained visual encoders, generating compact visual embeddings for key frames to build an efficient index. These visual embeddings, along with associated bounding boxes, are organized in an inverted multi-index structure within a vector database, which supports queries for any objects. During the query phase, LOVO transforms object queries to query embeddings and conducts fast approximate nearest-neighbor searches on the visual embeddings. Finally, a cross-modal rerank is performed to refine the results by fusing visual features with detailed textual features. Evaluation on real-world video datasets demonstrates that LOVO outperforms existing methods in handling complex queries, with near-optimal query accuracy and up to 85x lower search latency, while significantly reducing index construction costs. This system redefines the state-of-theart object query approaches in video analysis, setting a new benchmark for complex object queries with a novel, scalable, and efficient approach that excels in dynamic environments. Yuxin Liu 0007, Yuezhang Peng, Hefeng Zhou, Jiong Lou, Chentao Wu, Wei Zhao 0001, Jie Li 0002 |
ICDE | 7 |
| 2025 | FAI-CXL: An Efficient Hardware-Accelerated Fairness-Aware CXL Memory Pool Management with Fine-Grained Cacheline-Level InterleavingabstractCompute Express Link (CXL) enables scalable memory disaggregation, allowing multiple hosts to share a global memory pool. However, existing CXL pooling designs require manual static configuration, underutilize link bandwidth, and suffer from performance unfairness among hosts with diverse memory intensities. This paper presents FAI-CXL, a hardwareaccelerated CXL memory pool unified management architecture that integrates fine-grained, weighted interleaving and adaptive fairness-aware scheduling into the CXL switch. The proposed interleaving mechanism operates at cacheline granularity, improving both link and device bandwidth utilization, while the scheduling policy dynamically adjusts priorities to mitigate unfairness under imbalanced workloads. We implement FAI-CXL in a cycle-accurate ChampSim + Ramulator simulation framework with realistic CXL protocol modeling. Across diverse workload sets, FAI-CXL achieves up to 21 % IPC improvement and 17 % throughput gain compared to conventional pooling approaches, while ensuring fair performance across heterogeneous hosts. Yunfei Gu, Chentao Wu |
ICPADS | 4 |
| 2025 | Decision Shuffle: Efficient Pre-scheduling System for Push-based Shuffle in DAG Computing FrameworksabstractIn large-scale data-parallel analytics, shuffle operations often become performance bottlenecks due to network overhead from all-to-all data movement and disk I/O overhead from write/read of persistent intermediate data. Push-based shuffle is widely adopted to mitigate this overhead by enabling sequential I/O through early transmission and pre-merge. However, existing push-based-shuffle scheduling strategies based on single-shuffle-based workload prediction and task scheduling fails to account for hierarchical data dependencies in practical scenarios involving complex DAG workflows, leading to load imbalance and poor data locality. Chi Zhang 0005, Chentao Wu, Jie Li 0002, Minyi Guo, Liqiang Zhang 0010 |
ICPP | 3 |
| 2025 | Leveraging Peer-Informed Label Consistency for Robust Graph Neural Networks with Noisy LabelsabstractGraph Neural Networks (GNNs) excel in many applications but struggle when trained with noisy labels, especially as noise can propagate through the graph structure. Despite recent progress in developing robust GNNs, few methods exploit the intrinsic properties of graph data to filter out noise. In this paper, we introduce ProCon, a novel framework that identifies mislabeled nodes by measuring label consistency among semantically similar peers, which are determined by feature similarity and graph adjacency. Mislabeled nodes typically exhibit lower consistency with these peers, a signal we measure using pseudo-labels derived from representational prototypes. A Gaussian Mixture Model is fitted to the consistency distribution to identify clean samples, which refine prototype quality in an iterative feedback loop. Experiments on multiple datasets demonstrate that ProCon significantly outperforms state-of-the-art methods, effectively mitigating label noise and enhancing GNN robustness. Kailai Li 0002, Jiawei Sun 0001, Jiong Lou, Zhanbo Feng, Hefeng Zhou, Chentao Wu, Guangtao Xue, Wei Zhao 0001, Jie Li 0002 |
IJCAI | 6 |
| 2025 | An Effective Uncorrectable Memory Error Prediction Framework by Exploiting UPH Indicators in Production EnvironmentsabstractUCEs (Uncorrectable memory errors) pose significant challenges to cloud computing systems, often resulting in catastrophic failures and crashes. Researchers have explored prediction approaches to address this issue. Previous studies have provided insights into memory error prediction, focusing on memory module part numbers and relationships between error code data. However, these efforts face challenges due to insufficient data features and suboptimal optimization, especially in production environments where hardware/software sparing techniques are widely deployed, the UCE ratio is low, and long lead time is required. To address these issues, our study first collect a large amount of memory data from different vendors in Huawei's production environment, which has deployed hardware/software sparing techniques, to provide more general data. Second, we exploit new indicators termed UPH (Unique, Pinx, and History) from this data, which play a crucial role in predicting UCEs. UPH offers a more profound understanding of the factors contributing to UCEs and demonstrates higher precision and recall. Then, we integrate existing indicators and UPH into our prediction framework and demonstrate the significance of UPH through indicator importance assessments. We also optimize the framework by determining an optimal sampling window. In production environments with long lead time and low UCE ratio, we improve the framework by implementing noise reduction, self-history learning, and a new scenario-based model selection approach. Experimental results demonstrate 19 % - 27 % increase in UCE prediction recall with 4 %-11 % increase in precision under different scenarios, outperforming state-of-the-art methods in production environments. Xiaobo Zheng, Lisha Qin, Wen Xia, Chentao Wu, Yunfei Gu, Qicong Lin, Huifang Jiao, Rubing Huang |
IPDPS | 5 |
| 2025 | Towards Comprehensive Legal Document Analysis: A Multi-Round RAG ApproachabstractLegal document review is a time-consuming and highly specialized task, and the capabilities of intelligent legal review systems are limited and insufficient to complete detailed reviews. Traditional methods struggle with cross-references, dependencies, and context-dependent clauses. Our work introduces a multi-round RAG framework for legal document analysis, which iteratively refines queries and aggregates context to improve recall and understanding. Experiments on diverse contracts show a recall of 78.67%, outperforming baseline (57.33%) and single-round RAG (74.67%). Our analysis shows that iterative refinement effectively filters irrelevant results despite reduced precision. The multi-round approach halves missed cross-clause dependencies but reveals limitations in numerical consistency and obligation scope detection. These insights advance RAG for legal applications and provide a foundation for future work on scalable and accurate contract review. Wutong Zhang, Hefeng Zhou, Yunshen Li, Yuxin Liu 0007, Jiong Lou, Chentao Wu, Jie Li 0002 |
ICMR | 7 |
| 2025 | GD$^2$: Robust Graph Learning under Label Noise via Dual-View Prediction DiscrepancyabstractGraph Neural Networks (GNNs) achieve strong performance in node classification tasks but exhibit substantial performance degradation under label noise. Despite recent advances in noise-robust learning, a principled approach that exploits the node-neighbor interdependencies inherent in graph data for label noise detection remains underexplored. To address this gap, we propose GD$^2$, a noise-aware \underline{G}raph learning framework that detects label noise by leveraging \underline{D}ual-view prediction \underline{D}iscrepancies. The framework contrasts the \textit{ego-view}, constructed from node-specific features, with the \textit{structure-view}, derived through the aggregation of neighboring representations. The resulting discrepancy captures disruptions in semantic coherence between individual node representations and the structural context, enabling effective identification of mislabeled nodes. Building upon this insight, we further introduce a view-specific training strategy that enhances noise detection by amplifying prediction divergence through differentiated view-specific supervision. Extensive experiments on multiple datasets and noise settings demonstrate that \name~achieves superior performance over state-of-the-art baselines. Kailai Li 0002, Jiong Lou, Jiawei Sun 0001, Honghong Zeng, Chentao Wu, Yuan Luo 0003, Wei Zhao 0001, Shouguo Du, Jie Li 0002 |
NeurIPS | 6 |
| 2025 | Fast and Synchronous Crash Consistency with Metadata Write-Once File System
Yanqi Pan, Wen Xia, Xiangyu Zou, Zhenhua Li 0001, Chentao Wu |
OSDI | 7 |
| 2025 | A stochastic learning algorithm for multi-agent game in mobile network: A Cross-Silo federated learning perspective
Junzhe Liu, Zhaojiacheng Zhou, Shijing Yuan, Jiong Lou, Chentao Wu, Jie Li 0002 |
Comput. Networks | 6 |
| 2025 | Dynamic-EC: an efficient dynamic erasure coding method for permissioned blockchain systems
Mizhipeng Zhang, Chentao Wu, Jie Li 0002, Minyi Guo |
Frontiers Comput. Sci. | 2 |
| 2025 | Understanding and mitigating dimensional collapse of Graph Contrastive Learning: A non-maximum removal approach
Jiawei Sun 0001, Ruoxin Chen, Jie Li 0002, Yue Ding 0001, Chentao Wu, Zhi Liu 0002, Junchi Yan |
Neural Networks | 5 |
| 2025 | FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
Chi Zhang 0005, Yunfei Gu, Chentao Wu, Jie Li 0002, Xusheng Chen |
Proc. VLDB Endow. | 4 |
| 2025 | Adaptive Incentivize for Federated Learning With Cloud-Edge Collaboration Under Multi-Level Information SharingabstractFederated Learning with Cloud-Edge Collaboration (FL-CEC) has emerged as a cutting-edge paradigm in distributed learning. Efficient resource investment incentive mechanisms are crucial to encouraging clients in FL-CEC to contribute the necessary data and computational resources for training. However, existing studies are inadequate in meeting the incentive design requirements under multi-level information-sharing scenarios. Moreover, current works often rely on specific functional relationships between resource investment and global model accuracy. To bridge these gaps, this paper investigates the incentive problem for data and computational resource investment under multi-level information-sharing levels. We design a resource investment incentive mechanism based on a weighted potential game without depending on any specific functional relationship between data investment and model accuracy. Furthermore, we propose four algorithms to solve resource investment strategies for different levels of information sharing. The complexity and convergence rates of the proposed algorithms are thoroughly analyzed. Finally, we construct a simulation incentive platform on the Aliyun. Extensive evaluations demonstrate that the proposed scheme effectively enhances social welfare, and improves collaborative training accuracy and efficiency. Shijing Yuan, Beiyu Dong, Jie Li 0002, Song Guo 0001, Hongyang Chen 0001, Chentao Wu, Jie Wu 0001, Wei Zhao 0001 |
IEEE Trans. Computers | 6 |
| 2025 | Efficient Online Computing Offloading for Budget- Constrained Cloud-Edge Collaborative Video Streaming SystemsabstractCloud-Edge Collaborative Architecture (CEA) is a prominent framework that provides low-latency and energy-efficient solutions for video stream processing. In Cloud-Edge Collaborative Video Streaming Systems (CEAVS), efficient online offloading strategies for video tasks are crucial for enhancing user experience. However, most existing works overlook budget constraints, which limits their applicability in real-world scenarios constrained by finite resources. Moreover, they fail to adequately address the heterogeneity of video task redundancies, leading to suboptimal utilization of CEAVS's limited resources. To bridge these gaps, we propose an Efficient Online Computing framework for CEAVS (EOCA) that jointly optimizes accuracy, energy consumption, and latency performance through adaptive online offloading and redundancy compression, without requiring future task information. Technically, we formulate computing offloading and adaptive compression under budget constraints as a stochastic optimization problem that maximizes system satisfaction, defined as a weighted combination of accuracy, latency, and energy performance. We employ Lyapunov optimization to decouple the long-term budget constraint. We prove that the decoupled problem is a generalized ordinal potential game and propose algorithms based on generalized Benders decomposition (GBD) and the best response to obtain Nash equilibrium strategies for computing offloading and task compression. Finally, we analyze EOCA's performance bound, convergence rate, and worst-case performance guarantees. Evaluations demonstrate that EOCA effectively improves satisfaction while effectively balancing satisfaction and computational overhead. Shijing Yuan, Yuxin Liu 0007, Song Guo 0001, Jie Li 0002, Hongyang Chen 0001, Chentao Wu, Yang Yang 0001 |
IEEE Trans. Cloud Comput. | 6 |
| 2025 | Temporal Gradient Inversion Attacks With Robust OptimizationabstractFederated Learning (FL) has emerged as a promising approach for collaborative model training without sharing private data. However, privacy concerns regarding information exchanged during FL have received significant research attention.Gradient Inversion Attacks (GIAs)have been proposed to reconstruct the private data retained by local clients from the exchanged gradients. While recovering private data, the data dimensions and the model complexity increase, which thwart data reconstruction by GIAs. Existing methods adopt prior knowledge about private data to overcome those challenges. In this article, we first observe that GIAs with gradients from a single iteration fail to reconstruct private data due to insufficient dimensions of leaked gradients, complex model architectures, and invalid gradient information. We investigate a Temporal Gradient Inversion Attack with a Robust Optimization framework, called TGIAs-RO, which recovers private data without any prior knowledge by leveraging multiple temporal gradients. To eliminate the negative impacts of outliers, e.g., invalid gradients for collaborative optimization, robust statistics are proposed. Theoretical guarantees on the recovery performance and robustness of TGIAs-RO against invalid gradients are also provided. Extensive empirical results on MNIST, CIFAR10, ImageNet and Reuters 21578 datasets show that the proposed TGIAs-RO with 10 temporal gradients improves reconstruction performance compared to state-of-the-art methods, even for large batch sizes (up to 128), complex models like ResNet18, and large datasets like ImageNet (224× 224pixels). Furthermore, the proposed attack method inspires further exploration of privacy-preserving methods in the context of FL. Bowen Li 0013, Hanlin Gu, Ruoxin Chen, Jie Li 0002, Chentao Wu, Na Ruan, Xueming Si, Lixin Fan |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | ESFL: Accelerating Poisonous Model Detection in Privacy-Preserving Federated LearningabstractPrivacy-preserving federated learning (PPFL) is a promising secure distributed learning paradigm, which enables collaborative training of a global machine learning model through sharing encrypted local models instead of sensitive raw data. PPFL, however, is vulnerable to model poisoning attacks. Most existing Byzantine-robust PPFL solutions typically employ two non-colluding servers to achieve secure model detection and aggregation by executing interactive security protocols, which incur considerable computation and communication overheads. To tackle this issue, we propose an efficient and secure federated learning (ESFL) technique to accelerate the detection of poisonous models in PPFL. First, to improve computational efficiency, we construct a lightweight non-interactive efficient decryption functional encryption (NED-FE) scheme to protect the data privacy of local models. Then, to ensure high communication performance, we elaborately design a non-interactive privacy-preserving robust aggregation strategy, which efficiently detects the blind poisonous models and aggregates benign models. Finally, we implement ESFL and conduct extensive theoretical analysis and experiments. The numerical results demonstrate that ESFL not only achieves the confidentiality and robustness design goals but also maintains high efficiency. Compared with the baseline, ESFL effectively reduces the aggregation latency by up to 88%. Honghong Zeng, Jiong Lou, Kailai Li 0002, Chentao Wu, Guangtao Xue, Yuan Luo 0003, Fan Cheng 0002, Wei Zhao 0001, Jie Li 0002 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | V2PCP: Toward Online Booking Mechanism for Private Charging PilesabstractAs the adoption of electric vehicles continues to grow, the demand for extensive charging infrastructure in urban areas is concurrently rising. In response to the evolving charging infrastructure shortage, private charging piles have emerged as crucial supplementary energy sources, especially in areas lacking public charging infrastructure. The sharing of private charging piles, however, introduces several challenges. Notably, the variable availability time and extremely limited usage space of private charging piles pose scheduling complexities for charging pile owners. Furthermore, the completely peer-to-peer operation of private charging piles may lead to suboptimal solutions for fulfilling overall charging demand. To comprehensively address these challenges, we explore the potential for cooperation among geographically proximate charging piles. We introduce a novel online booking mechanism paired with specialized scheduling algorithms designed for scenarios involving both multiple private charging piles and single private charging piles. Our objective is to maximize the attained revenue of charging pile owners under fully dynamic conditions on both the supply and demand sides. Through meticulous theoretical proofs, we show that our mechanism achieves advantageous competitive ratios for both scenarios when compared to the offline optimal solutions. Numerous experiments, conducted with real charging sessions, consistently demonstrate that the proposed mechanism achieves the highest revenue, providing substantial evidence for its superior performance. Jiawei Sun 0001, Jiong Lou, Yusheng Ji, Chentao Wu, Wei Zhao 0001, Guangtao Xue, Yuan Luo 0003, Fan Cheng 0002, Jie Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Towards Bi-Level Supply/Demand Balanced Charging Systems via Online Power SchedulingabstractWith the rise of transportation electrification, an increasing number of charging stations have been established, forming a city-scale charging system. These charging stations serve as intermediaries that connectsupplyanddemand, drawing power from the grid and renewable energy sources to provide electricity to electric vehicles. Maintaining a delicate balance between supply and demand has emerged as a significant challenge for the charging system. On amacroscopiclevel, it impacts the power grid's peak load and reliability, whilelocally, it influences electric vehicle detour events. To comprehensively model the spatio-temporal characteristics in the charging system, we partition the charging system by adopting a supply-demand-aware approach and propose OPS, an online power scheduling algorithm based on the regularization technique. OPS aims to achieve a bi-level balance between supply and demand while constraining the power output of the charging system. We substantiate the efficacy of OPS through rigorous theoretical proofs, demonstrating its comparability to the optimal solution. Furthermore, we conduct extensive evaluation experiments with real-world data sets to establish the feasibility of the proposed methodology in alleviating the supply-demand imbalance. The results indicate that OPS attains an empirical competitive ratio of less than 1.2. Jiong Lou, Jie Li 0002, Runhui Xu, Chentao Wu, Zhi Liu 0002, Yuan Luo 0003, Yang Yang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Adaptive Incentive and Resource Allocation for Blockchain-Supported Edge Video Streaming Systems: A Cooperative Learning ApproachabstractEdge computing significantly enhanced the growth of edge-assistant video streaming applications. However, challenges such as unpredictable wireless conditions, resource constraints, and task redundancy have intertwined impacts on the overall performance of edge video streaming systems (EVS). Therefore, it is essential to have an integrated framework that addresses resource management, computational offloading, and video task preprocessing. Existing optimization strategies often neglect the simultaneous management of computational offloading, resource allocation, and video task preprocessing, leading to a suboptimal system utility. Moreover, they struggle to handle high-dimensional decision variables. On the other hand, learning-based adaptive schemes fall short in integrating distributed decisions and ensuring the scalability of wireless devices. Additionally, current approaches lack adaptive incentives. To bridge these gaps, we propose a novel framework called AIRA, which is based on improved multi-agent reinforcement learning (MARL) and smart contracts. AIRA manages resources, video compression, and adaptive incentives in a distributed manner. It consists of a MARL-driven cooperative learning algorithm (CLA) and a smart contract-guided adaptive incentive mechanism. Leveraging an actor-critic structure, the CLA enables wireless devices to master strategies for resource allocation, video task compression, and offloading, utilizing historical data. Notably, the CLA incorporates an attention mechanism to select pivotal tuples from the observation-action pairings among different agents, ensuring improved scalability and computational prowess. Evaluations based on real-world trajectories demonstrate that AIRA enables adaptive incentives. Compared to state-of-the-art approaches, CLA effectively enhances the long-term system utility and scalability of EVS. Shijing Yuan, Qingshi Zhou, Jie Li 0002, Song Guo 0001, Hongyang Chen 0001, Chentao Wu, Yang Yang 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | CAST: Cluster-Driven Truthful Crowdfunding Mechanism for Shared AI Service Deployment
Junzhe Liu, Shijing Yuan, Jiong Lou, Chentao Wu, Jie Li 0002 |
IEEE Trans. Serv. Comput. | 6 |
| 2024 | RL-Cache: An Efficient Reinforcement Learning Based Cache Partitioning Approach for Multi-Tenant CDN ServicesabstractContent Delivery Network (CDN) has been widely used to provide data transmission services to end users. The edge cache servers are important components in CDN, and their hit ratios significantly influence the quality of cache service. However, edge caches are shared by multiple tenants (i.e., Internet Content Providers or ICPs) and the resource contention among tenants presents a huge challenge to improve the cache performance. Cache partitioning is a common method to deal with this chal-lenge, and several approaches have been proposed but still have some drawbacks. Existing methods bring non-negligible temporal and spatial overheads while obtaining features. Although some learning based methods have reduced these costs, the learning model convergence is slow due to the large searching space. To address the above problems, we propose a lightweight Reinforcement Learning based Cache Partitioning Approach (RL-Cache), which increases overall hit ratios of edge cache servers in CDN. The core of RL-Cache is a new feature named Compulsory Miss Ratio (CMR). It can be obtained in linear complexity and reflect the tenants' demand of cache space. To demonstrate the effectiveness of our approach, we not only utilize open-source traces from industrial CDNs but also collect real-world workloads from Tencent Cloud CDN. We develop a simulator to conduct several experiments driven by various traces. The experimental results show that compared to the commonly used methods, RL-Cache reduces the upstream traffic by 12.6% on average and improves the hit ratio by up to 4%. Ranhao Jia, Zixiao Chen, Chentao Wu, Jie Li 0002, Minyi Guo, Hongwen Huang |
CLUSTER | 3 |
| 2024 | Efficient Serverless Function Scheduling in Edge ComputingabstractServerless computing is a promising approach for edge computing since its inherent features, e.g., lightweight virtualization, rapid scalability, and economic efficiency. However, there are two challenges existing in serverless edge computing: significant cold start latency and request blocking. Previous studies have not successfully resolved these challenges, which affect the Quality of Experience. In this paper, we formulate the Serverless Function Scheduling (SFS) problem in resource-limited edge computing, aiming to minimize the average response time. To solve this intractable scheduling problem, we first consider a simplified offline form of the SFS problem and design a polynomial-time optimal scheduling algorithm. Inspired by this optimal algorithm, we propose an Enhanced Shortest Function First (ESFF) algorithm, including function creation and function replacement. To avoid frequent cold starts, ESFF selectively decides the initialization of new function instances when receiving requests. To deal with request blocking, ESFF judiciously replaces serverless functions based on the function weight at the completion time of requests. Extensive simulations based on real-world serverless request traces are conducted, and the results show that ESFF consistently and substantially outperforms existing baselines under different settings. Jiong Lou, Zhiqing Tang, Shijing Yuan, Jie Li 0002, Weijia Jia 0001, Chentao Wu |
ICC | 7 |
| 2024 | Online Data Trading for Cloud-Edge Collaboration ArchitectureabstractCloud-edge collaboration Architecture (CEA) enables the co-training of AI models by cloud servers and edge servers, offering a promising solution for large-scale model training. An efficient data trading mechanism helps encourage edges to invest data resources to participate in training while reducing the cost of cloud servers. Existing research on data trading within CEA focuses on static scenarios, either overlooking the dynamics of data demand and the fairness of the selected edges or assuming unknown future communication overheads. To bridge these gaps and consider the long-term fairness constraints, we propose an Online Data Trading mechanism for the CEA, called ODT, to improve the long-term utility. Technically, ODT decouples the long-term fairness constraint into a series of single time-slot sub-problems using the Lyapunov optimization method and applies dynamic programming to solve the single time-slot edge selection sub-problems. We prove the NP-hardness of the sub-problems, the performance bounds, and the computational complexity of the proposed algorithm. Evaluation results demonstrate that the proposed mechanism effectively improves long-term utility and achieves an efficient trade-off between fairness and utility. Shijing Yuan, Jie Li 0002, Jiong Lou, Chentao Wu, Song Guo 0001, Yang Yang 0001 |
ICC | 5 |
| 2024 | GCC: Optimizing Space Efficiency and Read Latency of SSDs with Workload-Aware Garbage Collection Aided CompressionabstractData compression is increasingly employed to enhance throughput and space efficiency in flash-based storage systems, which are critical for data-intensive applications. Current intra-SSD compression techniques operate transparently with respect to the file system and contribute to improving the lifetime of SSDs. These approaches typically avoid compressing read-hot data to reduce the latency penalties associated with decompression. However, the read-hot data remain uncompressed even after turning into cold data, thereby reducing overall compression effectiveness and diminishing space efficiency. Moreover, when previously compressed cold data become read-hot, it necessitates frequent decompression, which increases the read latency. To address the above problems, we propose a novel Garbage Collection aided Compression (GCC) scheme, to optimize space efficiency and mitigate read latency for compression-supported SSDs. The key idea of GCC is exploiting the valid page migration during garbage collection to enable background compression and decompression. Throughout the garbage collection process, the migrated valid pages can potentially be compressed or decompressed, which progressively improves space efficiency and minimizes the need for decompression during read operations. Performance evaluations conducted using MQSim simulator demonstrate that, compared to the typical compression schemes, GCC reduces the read and write latency by 25.27% and 9.43% on average and improves the space efficiency by 15.01% on average. Linhui Liu, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo |
ICCD | 4 |
| 2024 | HGR: A Hybrid Global Graph-Based Recovery Approach for Cloud Storage Systems with Failure and Straggler NodesabstractCloud storage systems often face the issues of failure and straggler nodes. Failure is characterized as a fail-stop scenario, which refers to disk failures that can result in significant data unavailability. Straggler nodes are typically those with heavy workloads or poor performance. Usually, both failure and straggler nodes coexist, posing a significant challenge to data availability in storage systems. In such failure scenarios, parallel recovery and straggler recovery methods are commonly used as separate approaches for data recovery. However, parallel recovery methods encounter bottlenecks on the recovery path due to the presence of straggler nodes. Meanwhile, straggler recovery methods face the challenge of lacking available recovery paths in cases of multiple node failures. Scenarios involving both multiple failures and stragglers are common, yet there is a lack of efficient recovery methods for these situations. In this paper, we focus on scenarios involving video data, which occupies a significant portion of cloud storage systems, to address the above issues. We propose a Hybrid Global Graph-based Recovery (HGR) method that integrates parallel and straggler recovery approaches into a single global graph. The key idea of HGR is to construct a global graph that includes global node parameter information, enabling comprehensive coordination. We partition the global graph into two subgraphs: one containing straggler nodes and the other containing failure nodes. Resources are efficiently allocated to each subgraph to schedule recovery tasks in parallel. For data that presents significant recovery challenges, exhibits poor parallelism, has substantial tail latency, or exceeds fault tolerance limits, we employ approximate recovery methods. To demonstrate HGR's effectiveness, we conducted several experiments. The results indicate that HGR can reduce recovery time by up to 45.06% and improve I/O throughput by as much as 1.79× compared to state-of-the-art recovery methods. Piao Hu, Huangzhen Xue, Chentao Wu, Minyi Guo, Jie Li 0002, Xiangyu Chen 0006, Shaoteng Liu, Liyang Zhou, Shenghong Xie |
ICDCS | 3 |
| 2024 | InterpGNN: Understand and Improve Generalization Ability of Transdutive GNNs through the Lens of Interplay between Train and Test NodesabstractTransductive node prediction has been a popular learning setting in Graph Neural Networks (GNNs). It has been widely observed that the shortage of information flow between the distant nodes and intra-batch nodes (for large-scale graphs) often hurt the generalization of GNNs which overwhelmingly adopt message-passing. Yet there is still no formal and direct theoretical results to quantitatively capture the underlying mechanism, despite the recent advance in both theoretical and empirical studies for GNN's generalization ability. In this paper, the $L$-hop interplay (i.e., message passing capability with training nodes) for a $L$-layer GNN is successfully incorporated in our derived PAC-Bayesian bound for GNNs in the semi-supervised transductive setting. In other words, we quantitatively show how the interplay between training and testing sets influence the generalization ability which also partly explains the effectiveness of some existing empirical methods for enhancing generalization. Based on this result, we further design a plug-and-play ***Graph** **G**lobal **W**orkspace* module for GNNs (InterpGNN-GW) to enhance the interplay, utilizing the key-value attention mechanism to summarize crucial nodes' embeddings into memory and broadcast the memory to all nodes, in contrast to the pairwise attention scheme in previous graph transformers. Extensive experiments on both small-scale and large-scale graph datasets validate the effectiveness of our theory and approaches. Jiawei Sun 0001, Kailai Li 0002, Ruoxin Chen, Jie Li 0002, Chentao Wu, Yue Ding 0001, Junchi Yan |
ICLR | 5 |
| 2024 | HMT: A Hybrid Mitigating and Transferring Approach on I/O Throughput Degradation for Erasure Coded Storage SystemsabstractIn cloud storage systems with erasure coding (EC), increased demand for data services and EC-based data recovery lead to high volumes of concurrent I/O requests, potentially causing network congestion or server overload. Network congestion or node overload significantly reduces I/O throughput and data parallelism. Various methods have been proposed to address these issues, such as fine-grained data packet partitioning, I/O scheduling, and transfer reading. However, these methods may not be effective in different scenarios. For instance, even if most I/O paths are relieved, data may still remain inaccessible. Piao Hu, Huangzhen Xue, Chentao Wu, Jie Li 0002, Minyi Guo |
ICPP | 3 |
| 2024 | SecureCut: Federated Gradient Boosting Decision Trees with Efficient Machine Unlearning
Bowen Li 0013, Jie Li 0002, Chentao Wu |
ICPR (5) | 4 |
| 2024 | CKSM: An Efficient Memory Deduplication Method for Container-based Cloud Computing SystemsabstractMemory deduplication techniques are widely used to improve memory utilization in cloud computing platforms, and they can be categorized into virtualized and containerized environments. In virtualized environments, prevalent memory deduplication approaches often rely on scanning the virtual address space of different processes. However, the complexity of virtual address spaces can reduce scanning efficiency in containerized environments. Additionally, the many-to-one mapping between virtual and physical pages can decrease the efficiency of merging operations.To solve the above problems, we proposed a Container-based Kernel Samepage Merging method called CKSM. This method leverages potential duplicate candidates and efficiently performs merging operations. It employs layered sampling to construct the priority of physical pages. Additionally, a physical page scanning mechanism is designed to directly obtain valid pages within the system. CKSM uses the physical page merge mechanism to merge all virtual pages at once and release the corresponding memory directly. We conduct several experiments to demonstrate the efficiency of CKSM. It reduces the scanning overhead by up to 80.99% and increases page comparison efficiency by up to 42.51%. Besides, CKSM achieves an average of 3.02×memory usage reduction compared to UKSM and 2.79×response speedup compared to KSM in the containerized environment. In cloud computing emulation, CKSM has been proven to be optimal in high-density deployment. Yunfei Gu, Yihui Lu, Chentao Wu, Jie Li 0002, Minyi Guo |
IPDPS | 3 |
| 2024 | A Parallel Partial Merge Repair Algorithm for Multi-block Failures for Erasure Storage SystemsabstractIn order to achieve high availability and low storage costs in distributed storage systems, erasure code is widely used instead of replication. Compared to replication, erasure code can reduce storage costs, but also brings higher repair costs. There are currently many repair algorithms to reduce the block reconstruction time of single block failure. However, applying the existing methods to multi-block failures may lead to unbalanced network traffic, unnecessary network transfers, and network congestion at data collection node during the repair process, which can not make full use of the bandwidth between nodes.To solve this problem, we propose a novel repair algorithm called Partial Merge Repair (PMR) for multi-block failures, which is a scheduling algorithm that considers network load between nodes and combines multiple failed blocks to recover together. It first divides all surviving nodes into different groups, and then the data collection nodes within the group collect the data needed to repair multiple blocks through cross merging. Finally, the data collection node sends the collected blocks to the repair node to complete the repair. Our study presents a formal definition and proof of network transfer time in the modeled repair process of PMR, highlighting its superior efficiency compared to existing methods in homogeneous environments.We implement a prototype of PMR to evaluate its performance. The experimental results indicate that compared to existing repair technologies, PMR improves repair throughput by 28%-256% for various scenes. Shuaipeng Zhang, Chentao Wu, Ruobin Wu, Saiqin Long, Wen Xia |
IPDPS | 3 |
| 2024 | Turbo Table: A Semantic-Aware Cache Acceleration SystemabstractIn the current storage disaggregation architecture, the challenge of quickly retrieving data from storage clusters is typically addressed using caching or data pushdown strategies to accelerate data access and reduce data movement. However, most caching systems still manage data at the page granularity level. For compute-intensive applications, such as analytical applications, this coarse data management approach leads to underutilized computational resources and read-write amplification issues. Additionally, insufficient cache utilization results in inefficient data flow. By managing data at the schema granularity level and modifying the data flow path, we alleviate the read-write amplification problem and improve data transfer speeds. Managing data at the schema level also enables semantic awareness, allowing us to proactively analyze semantics for more precise cache management strategies and accurate prefetching, rather than passively waiting for cache request sequences. We also observed that native compute caches waste valuable cache space and complicate the association between original and result data. To address this, we propose a multi-grained caching model to avoid these limitations. Compared to the baseline, Turbo Table reduces computation time by 2.4% to 36.8%. Chentao Wu, Jie Li 0002, Minyi Guo |
ISPA | 2 |
| 2024 | Exploit both SMART Attributes and NAND Flash Wear Characteristics to Effectively Forecast SSD-based Storage Failures in Clusters
Yunfei Gu, Chentao Wu, Xubin He |
USENIX ATC | 2 |
| 2024 | Adaptive Incentive for Cross-Silo Federated Learning in IIoT: A Multiagent Reinforcement Learning ApproachabstractIn the Industrial Internet of Things (IIoT), cross-silo federated learning (CSFL) enables entities, such as manufacturers and suppliers to train global models for optimizing production processes while ensuring data privacy. A well-designed incentive mechanism is essential to persuade clients to contribute data resources. However, existing methodologies overlook the dynamic nature of the training process, where the accuracy of the globally trained model and the client’s data ownership change over time. Furthermore, the majority of previous research assumes a defined functional relationship between the data contribution and the model accuracy, which is infeasible in realistic and dynamic training environments. To address these challenges, we design a novel adaptive mechanism for CSFL that inspires organizations to contribute data resources in a dynamic training environment with the aim of maximizing their long-term payoffs. This mechanism leverages multiagent reinforcement learning (MARL) to ascertain near-optimal data contribution strategies from potential game histories without necessitating private organizational information or a precise accuracy function. Experimental results indicate that our mechanism achieves adaptive incentive in dynamic environments and effectively enhances the long-term payoffs of organizations. Shijing Yuan, Beiyu Dong, Hongtao Lv, Hongyang Chen 0001, Chentao Wu, Song Guo 0001, Yue Ding 0001, Jie Li 0002 |
IEEE Internet Things J. | 6 |
| 2024 | Ada-WL: An Adaptive Wear-Leveling Aware Data Migration Approach for Flexible SSD Array Scaling in ClustersabstractRecently, the flash-based Solid State Drive (SSD) array has been widely implemented in real-world large-scale clusters. With the increasing number of users in upper-tier applications and the burst of Input/Output requests in this data explosive era, data centers need to continuously scale up to meet real-time data storage needs. However, the classical disk array scaling methods are designed based on HDDs, ignoring the wear leveling and garbage collection characteristics of SSD. This leads to penalties due to the vast lifetime gap between extended SSDs and the original in-use SSDs while scaling the SSD array, including extra triggered wear leveling I/O, latency in average response time, etc.To address these problems, we propose an Adaptive Wear-Leveling aware data migration approach for flexible SSD array scaling in clusters. It manages the interdisk wear leveling based on Model Reference Adaptive Control, which includes an SSD behavior emulator, Kalman filter estimator, and adaptive law. To demonstrate the effectiveness of this approach, we conducted several simulations and implementations on actual hardware. The evaluation results show that Ada-WL has the self-adaptability to optimize the wear leveling management parameters for various states of SSD arrays, diverse workloads, and scaling performed multiple times, significantly improving performance for SSD array scaling. Yunfei Gu, Linhui Liu, Chentao Wu, Jie Li 0002, Minyi Guo |
IEEE Trans. Computers | 3 |
| 2024 | BSR-FL: An Efficient Byzantine-Robust Privacy-Preserving Federated Learning FrameworkabstractFederated learning (FL) is a technique that enables clients to collaboratively train a model by sharing local models instead of raw private data. However, existing reconstruction attacks can recover the sensitive training samples from the shared models. Additionally, the emerging poisoning attacks also pose severe threats to the security of FL. However, most existing Byzantine-robust privacy-preserving federated learning solutions either reduce the accuracy of aggregated models or introduce significant computation and communication overheads. In this paper, we propose a novelBlockchain-basedSecure andRobustFederatedLearning (BSR-FL) framework to mitigate reconstruction attacks and poisoning attacks. BSR-FL avoids accuracy loss while ensuring efficient privacy protection and Byzantine robustness. Specifically, we first construct a lightweight non-interactive functional encryption (NIFE) scheme to protect the privacy of local models while maintaining high communication performance. Then, we propose a privacy-preserving defensive aggregation strategy based on NIFE, which can resist encrypted poisoning attacks without compromising model privacy through secure cosine similarity and incentive-based Byzantine-tolerance aggregation. Finally, we utilize the blockchain system to assist in facilitating the processes of federated learning and the implementation of protocols. Extensive theoretical analysis and experiments demonstrate that our new BSR-FL has enhanced privacy security, robustness, and high efficiency. Honghong Zeng, Jie Li 0002, Jiong Lou, Shijing Yuan, Chentao Wu, Wei Zhao 0001, Sijin Wu |
IEEE Trans. Computers | 5 |
| 2024 | Toward Real-Time Pricing and Allocation for Surplus Resources in Electric Bus Charging StationsabstractWe are witnessing a rapid growth of electric vehicles in both individual and public transportation. Many large-scale company-built electric bus charging stations have been established to facilitate public transportation, while these electric bus charging stations are not accessible to private vehicles. The motivation for this work comes from the possibility of opening surplus resources in electric bus charging stations to alleviate the charging resource shortage and gain extra profit for the electric bus charging station. The operation facing private vehicles, however, also encounters obstacles including serious congestion induced by absorbing private vehicles and delays of bus lines due to private vehicles occupying charging points. To jointly solve these challenges, we propose a real-time control mechanism to maximize the long-term net profit of an electric bus charging station based on Lyapunov optimization theory and generalized benders decomposition, while maintaining the congestion level and ensuring timetables of buses. We demonstrate through rigorous theoretical proof that the proposed mechanism can be arbitrarily close to the optimal solution. Comprehensive evaluation experiments with real-world data sets have been conducted to show the credibility of the mechanism in reducing congestion, ensuring bus timetables, and maximizing the long-term net profit. Jie Li 0002, Shijing Yuan, Haiming Jin, Chentao Wu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsabstractAlong with the fast evolution of deep neural networks, the hardware system is also developing rapidly. As a promising solution achieving high scalability and low manufacturing cost, multi-accelerator systems widely exist in data centers, cloud platforms, and SoCs. Thus, a challenging problem arises in multi-accelerator systems: selecting a proper combination of accelerators from available designs and searching for efficient DNN mapping strategies. To this end, we propose MARS, a novel mapping framework that can perform computation-aware accelerator selection, and apply communication-aware sharding strategies to maximize parallelism. Experimental results show that MARS can achieve 32.2% latency reduction on average for typical DNN workloads compared to the baseline, and 59.4% latency reduction on heterogeneous models compared to the corresponding state-of-the-art method. Guan Shen, Jieru Zhao, Zeke Wang, Zhe Lin 0007, Wenchao Ding 0001, Chentao Wu, Quan Chen 0002, Minyi Guo |
DAC | 6 |
| 2023 | Adaptive Processing for Video Streaming with Energy Constraint: A Multi-Agent Reinforcement Learning MethodabstractEdge computing is a highly promising technology that empowers mobile devices to offload video streaming tasks to edge servers, thereby improving the video stream analysis performance. However, most existing research on edge video streaming has failed to give adequate attention to the joint optimization of video streaming tasks with respect to dynamics, redundancy, and long-term energy constraints. To address this limitation, we propose a novel method based on a multi-agent reinforcement learning algorithm, which significantly enhances the performance of edge video stream analysis under long-term energy constraints. Specifically, our proposed method conducts video compression and offloading under long-term energy constraints to maximize the long-term rewards of video task processing. Experimental evaluations have demonstrated the convergence of the proposed method, which outperforms the baseline solutions, achieving higher long-term rewards. Haotian Fu, Shijing Yuan, Chentao Wu, Yuan Luo 0003, Jie Li 0002 |
GLOBECOM | 4 |
| 2023 | Towards Practical Edge Inference Attacks Against Graph Neural NetworksabstractGraph Neural Networks (GNNs) have demonstrated superior performance in numerous real-world applications. Despite their success, recent studies have shown that GNNs are vulnerable under edge inference attacks aimed to infer the connectivity of a given pair of nodes. However, existing methods primarily focus on the scenario when properties of target nodes are revealed. In this paper, we propose an edge inference attack in a more realistic and practical setting. In our threat model, the adversary cannot obtain properties of target nodes but can inject a single probing node and query the target GNN for its prediction. By connecting the probing and target nodes, the adversary can infer the connectivity of the target node pair based on the prediction of the probing node. Extensive experiments show that our attack performs comparably to ones that require properties of target nodes. And when given such auxiliary knowledge, our attack outperforms state-of-the-art methods. Kailai Li 0002, Jiawei Sun 0001, Ruoxin Chen, Kexue Yu, Jie Li 0002, Chentao Wu |
ICASSP | 7 |
| 2023 | TradeFL: A Trading Mechanism for Cross-Silo Federated LearningabstractCross-silo federated learning (CFL) is a distributed learning paradigm that allows organizations (e.g., financial or medical entities) to train a global model on siloed data. Recent studies on mechanisms designed for CFL, however, rarely jointly consider the potential inter-organizational competition and the lack of credibility between organizations, which may discourage organizational participation. In this paper, we investigate the problem of inter-organizational competition and credibility assurance. We propose a distributed trading mechanism, called$TradeFL$, to incentivize organizations to contribute data and computational resources through mutual trading among organizations. Technically, TradeFL characterizes the competition among organizations and compensates for their damage incurred by competition. TradeFL runs on distributed organizations and provides credibility guarantees for compensation through a customized smart contract11Illustration of the prototype: https://github.com/user10963.. We prove that the interaction among organizations that contribute resources to maximize personal payoffs is a weighted potential game. Then, we propose a centralized algorithm and a distributed algorithm to determine the optimal resource contribution. Simulation results and evaluations based on real-world datasets demonstrate that our scheme achieves higher social welfare, increases the amount of contributed data by up to 64%, and improves the accuracy of the global model by at most 23.2%. Shijing Yuan, Hongtao Lv, Chentao Wu, Song Guo 0001, Zhi Liu 0002, Hongyang Chen 0001, Jie Li 0002 |
ICDCS | 4 |
| 2023 | DW-LRC: A Dynamic Wide-stripe LRC Codes for Blockchain Data Under Malicious Node ScenariosabstractBlockchain is a decentralized digital ledger system that can be used in many fields. However, the traditional approach of storing blockchain data (full repliction) incurs expensive storage costs. As a result, researchers have proposed erasure coding methods to reduce storage costs. Existing erasure coding methods are all based on RS codes. Since blockchain systems typically have a large number of nodes, it is necessary to use wide-stripe RS codes to store blocks, which results in significant overhead during the recovery processes. Because widestripe RS codes require accessing a large number of nodes for decoding, this means significant network and I/O cost.To solve the above problem, we propose DW-LRC, which is a dynamic wide-stripe Local Reconstruction Code (LRC) based methods in permissioned blockchain systems. DW-LRC predicts malicious nodes using node reputation, and then selects an appropriate variant of LRC code to lower storage and recovery costs compared to traditional erasure coding. To demonstrate the effectiveness of DW-LRC, we conduct several experiments on the open source blockchain software tendermint. The results show that, compared to the state-of-the-art erasure coding methods, DW-LRC reduced the average recovery latency by 36.4% and improved the block access throughput by 32.3%. Mizhipeng Zhang, Chentao Wu, Jie Li 0002, Minyi Guo |
ICPADS | 2 |
| 2023 | Inductive Dummy-based Homogeneous Neighborhood Augmentation for Graph Collaborative FilteringabstractIn the era of information explosion, we urgently need recommendation systems to filter massive amounts of information. Recent advancements in graph neural networks have led to the widespread adoption of graph collaborative filtering algorithms for recommendation systems. Despite their effectiveness, graph collaborative filtering algorithms have several limitations, such as data sparsity and long-tailed distribution. This sparse data with numerous long-tailed nodes can be viewed as an inductive scenario in which models require a robust inductive ability to learn quality representations from sparse data. Inductive graph collaborative filtering methods, such as Pin-SAGE, improve the generalization ability via random neighbor sampling. However, these inductive methods are time-consuming or ineffective in transductive scenarios because of the complicated operations and information loss in random neighbor sampling methods. We propose IDHA, an inductive dummy-based homogeneous neighborhood augmentation method for graph collaborative filtering, to address the issues above. Our method employs dummy nodes connected to all nodes to take advantage of low-degree nodes in graph structure learning. To improve the model's generalization in inductive scenarios, we adopt one-hop random-walk sampling. We propose homogeneous neighborhood augmentation via contrastive learning for inductive graph collaborative filtering. This method exploits contrastive learning in sparse data to its full potential. In addition, we employ a lightweight model design to enhance performance and practicality while reducing model complexity. Extensive experiments on three datasets demonstrate that our method outperforms the state-of-the-art transductive and inductive graph collaborative filtering recommendation methods. Jiawei Sun 0001, Jie Li 0002, Chentao Wu |
IJCNN | 4 |
| 2023 | Improving Productivity and Efficiency of SSD Manufacturing Self-Test Process by Learning-Based Proactive Defect PredictionabstractIn the recent storage market, Flash-based Solid State Drives (SSDs) have become high-performance alternatives to Hard Disk Drives (HDDs), dramatically increasing SSD shipments. To guarantee product reliability and quality to remain competitive, SSD manufacturers pay significant efforts in technology qualification and reliability design, especially in Manufacturing Self-Test (MST) processes. However, the cost of the MST process becomes more prominent as the memory density of SSD increases. In this paper, we study the MST data in over 20,000 SSDs and propose a novel and economical approach to dynamically reduce the MST overhead by proactive infant defect prediction based on Generative Adversarial Network-Attention based Spatial-Temporal Sequence-to-Sequence network (GAN-ASTSeq). It reduces the temporal cost by 80.2% (i.e., improves the efficiency by 4×) while maintaining an outstanding detection rate of defects. Yunfei Gu, Zixiao Chen, Chentao Wu, Xinfei Guo, Jie Li 0002, Minyi Guo, Rong Yuan, Taile Zhang, Haoran Cai |
ITC | 4 |
| 2023 | JIRA: Joint Incentive Design and Resource Allocation for Edge-Based Real-Time Video Streaming SystemsabstractEdge computing has been introduced as a promising technology for real-time video streaming systems. However, due to the lack of automatic incentives and the limitation of resources, traditional edge computing performs poorly in nowadays scenarios. To handle these two challenges, we propose a framework ofJointIncentive design andResourceAllocation (JIRA) for edge-based real-time video streaming systems. Technically, to ensure the trust and automatic distribution of incentives, we develop a novel smart contract based incentive mechanism and implement a prototype. Meanwhile, we propose an efficient online algorithm, i.e., JIRA, which dynamically adjusts compression ratio, offloading decision, and resource allocation to achieve performance optimization for video streaming under long-term latency and resource constraints. Specifically, JIRA is based on Lyapunov optimization, which decomposes the challenging long-term decision problem into a series of real-time optimization problems. Then we propose a multi-cut Generalized Benders Decomposition based algorithm (MGA) to tackle the non-convexity of the decomposed problem. Through rigorous theoretical analysis, we prove the performance bound of JIRA. Extensive simulations demonstrate that the proposed schemes can achieve an efficient trade-off between accuracy performance and energy consumption. Shijing Yuan, Jie Li 0002, Hongyang Chen 0001, Zhu Han 0001, Chentao Wu, Yongbing Zhang 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2022 | An Energy-efficient Computing Offloading Framework for Blockchain-enabled Video Streaming SystemsabstractBlockchain and edge computing have been widely applied in video streaming systems. However, previous works lack a joint consideration of video redundancy and full utilization of edge resources (bandwidth resources, CPU frequency), resulting in suboptimal performance of video streaming systems. In this paper, we propose a computing offloading framework for blockchain-enabled video streaming systems to fully exploit edge resources and reduce energy consumption. Specifically, we formulate computing offloading, resource allocation, and adaptive compression as a joint optimization problem. We transform and decompose the original non-convex problem and propose an algorithm based on the alternating direction method of multipliers (ADMM) to solve the decomposed problem in a distributed manner. Simulation results demonstrate that our scheme can effectively reduce energy consumption and fully utilize the bandwidth and computational resources. Shijing Yuan, Jie Li 0002, Yuxuan Zhu 0003, Chentao Wu, Yue Ding 0001 |
GLOBECOM | 4 |
| 2022 | Ada-STNet: A Dynamic AdaBoost Spatio-Temporal Network for Traffic Flow PredictionabstractTraffic flow prediction is of particular interest since its massive applications in intelligent transportation systems (ITS). The problem is challenging due to the complex spatio-temporal correlations and nonlinearities of traffic flows. However, existing methods based on the graph neural networks cannot efficiently extract the dynamic and long-range spatial correlations, thus producing unsatisfactory prediction results. In this paper, we propose an AdaBoost Spatio-temporal Network (Ada-STNet). Similar to AdaBoost, Ada-STNet stacks several base neural networks as "layers" which capture spatial and temporal correlations simultaneously. Each layer learns an adaptive adjacency matrix from weights and embedding of nodes. The adjacency matrix is layer-wise adjusted to extract information from distant neighbors and adapt to dynamic correlations. Experiments are conducted on three real-world benchmark datasets, demonstrating that the Ada-STNet outperforms the state-of-the-art methods. Jiawei Sun 0001, Jie Li 0002, Chentao Wu, Zili Tang, Celimuge Wu |
ICASSP | 3 |
| 2022 | Iterative Learning for Distorted Image RestorationabstractDeep generative networks have achieved great success on distorted image restoration. However, existing deep learning approaches mainly focus on delicate module structure while ignoring the saturation problem. In this paper, we study the influence of different learning schemes on fitting capability and tackle the problem by proposing a novel iterative learning scheme. It accumulates weight importance from past episodes and guides the network to search for the optimal of current episodes based on obtained knowledge. Since public available datasets contain very few distortion types, we also release a new benchmark to explore this task. Extensive experimental evaluations on the benchmarks demonstrate that our learning approach significantly outperforms all other methods and achieves new state-of-the-art results. Chao Wang 0009, Jie Li 0002, Xinlei He 0006, Chentao Wu |
ICASSP | 7 |
| 2022 | RCS: A Redirection Computational Scheduler to Accelerate Straggler Recovery for Erasure Coded Cloud Storage SystemabstractThe straggler problem is one of the most significant problems in cloud computing systems, in which a large number of parallel processes are blocked by a small set of straggler tasks with a long waiting time. This problem is crucial in erasure coded storage systems, where the recovery processes require to retrieve a set of multiple chunks among different nodes. With skewed data accesses from various applications, several nodes with a high workload could easily become stragglers during the recovery process, leading to unacceptable long tail latency. To address the above problems, we propose a Redirection Computational Scheduling method called RCS, to accelerate the data recovery under straggler scenarios. The key idea of RCS is transferring the computational and network workload from one node to another, which can avoid the adverse effects caused by the stragglers. To demonstrate the effectiveness of RCS, we conduct several experiments in a cluster. The results show that, compared to the state-of-the-art recovery methods, RCS saves the recovery time by up to 72.1%, and speeds up the recovery throughput by up to a factor of 1.4X, respectively. Xinzhe Cao, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo, Yuanyuan Dong 0002 |
ICCD | 3 |
| 2022 | GRPU: An Efficient Graph-based Cross-Rack Parallel Update Scheme for Cloud Storage SystemsabstractErasure coding (EC) has been widely used in cloud storage systems to provide both high reliability and low storage cost. Previous literatures show that the cross-rack update operations are prevalent for many applications in erasure-coded cloud storage systems, which introduces significant I/O amplification, load imbalance and high latency. Several existing methods have been proposed to mitigate these problems. However, they ignore the correlations among chunks when performing data placement. Thus numerous stripes and racks participate in the update leading to extra I/Os and cross-rack traffic. Moreover, they don’t take into account the parallelism of network transmission which loses the potential update performance gains.To address the issues, we propose a novel Graph-based cross-Rack Parallel Update (GRPU) scheme to improve the update performance for erasure-coded cloud storage systems. The key idea of GRPU is to place the correlated chunks in the same stripe and rack, and transmit the chunks in parallel based on the network distance. The data placement and transmission paths selection are guided by two kinds of graphs. To demonstrate the effectiveness of GRPU, we conduct several experiments in a local cluster. The results show that, compared to the state-of-the-art methods, GRPU reduces the cross-rack traffic by up to 34.66% and the average response time by up to 61.69%, respectively. Ranhao Jia, Haiwei Deng, Yunfei Gu, Huangzhen Xue, Chentao Wu, Jie Li 0002, Guangtao Xue, Minyi Guo |
ICCD | 5 |
| 2022 | CRAB: Certified Patch Robustness Against Poisoning-Based Backdoor AttacksabstractBackdoor attacks have been proved to be seriously threatening to deep neural networks. Many defending methods against backdoor attack have been proposed and reduced attack success rate significantly. However, most existing defending methods are empirical, and might be later broken by stronger attack methods. To avoid such a cat-and-mouse game, We proposed CRAB, a defense that can guarantee the robustness of an image classifier against poisoning-based backdoor attack with triggers bounded in a contiguous region. We analyze two ways of adding triggers: fixed-region and randomized-region. For fixed-region setting, we train a set of models on benign dataset for different image ablation positions and give robustness guarantee to both training and testing datasets. Whilst for random position triggers, we train a universal model on dataset with triggers, and give robustness guarantee to testing datasets. Our excellent experimental results demonstrate that CRAB exhibits strong robustness against patched backdoor attack, while maintaining comparable high clean accuracies. Huxiao Ji, Jie Li 0002, Chentao Wu |
ICIP | 3 |
| 2022 | Zero-Shot Scene Graph Generation with Knowledge Graph CompletionabstractLimited by the incomprehensive training samples, existing scene graph generation (SGG) methods perform poorly on predicting zero-shot (i.e., unseen) subject-predicate-object triples. To address this problem, we propose a general SGG framework to improve their zero-shot performance. The main idea of our method is to generate the information of zero-shot triples before the training of the predicate classifier and thus make the original zero-shot triples non-zero-shot. Specifically, the missing information of zero-shot triples is generated by our proposed knowledge graph completion strategy and then integrated with visual features of images. Therefore, the predicate classification of zero-shot triples is no longer just regarded as a single visual classification task but also transformed into a prediction task of missing links in a knowledge graph. The experiments on the dataset Visual Genome demonstrate that our proposed method outperforms the state-of-the-art methods in popular zero-shot metrics (i.e., zR@N, ng-zR@N) for all popular SGG tasks. Ruoxin Chen, Jie Li 0002, Jiawei Sun 0001, Shijing Yuan, Huxiao Ji, Chentao Wu |
ICME | 8 |
| 2022 | On Collective Robustness of Bagging Against Data PoisoningabstractBootstrap aggregating (bagging) is an effective ensemble protocol, which is believed can enhance robustness by its majority voting mechanism. Recent works further prove the sample-wise robustness certificates for certain forms of bagging (e.g. partition aggregation). Beyond these particular forms, in this paper, we propose the first collective certification for general bagging to compute the tight robustness against the global poisoning attack. Specifically, we compute the maximum number of simultaneously changed predictions via solving a binary integer linear programming (BILP) problem. Then we analyze the robustness of vanilla bagging and give the upper bound of the tolerable poison budget. Based on this analysis, we propose hash bagging to improve the robustness of vanilla bagging almost for free. This is achieved by modifying the random subsampling in vanilla bagging to a hash-based deterministic subsampling, as a way of controlling the influence scope for each poisoning sample universally. Our extensive experiments show the notable advantage in terms of applicability and robustness. Our code is available at https://github.com/Emiyalzn/ICML22-CRB. Ruoxin Chen, Zenan Li, Jie Li 0002, Junchi Yan, Chentao Wu |
ICML | 5 |
| 2022 | ERP: An Efficient Rewrite Scheme to Improve the Inline Deduplication Restore Performance in Backup SystemsabstractData deduplication is an effective technique to reduce the amount of redundant data, which is widely used in backup systems. To reconstruct the original backup data, restore is a typical procedure in inline deduplication, which brings read amplification when duplicate chunks are shared among various data streams. Rewrite is a cost-efficient method to improve the inline deduplication restore performance by writing the fragmented duplicate chunks repeatedly. Although several rewrite schemes are proposed to improve the restore performance, they either decrease the deduplication ratio or increase the temporal overhead of the rewrite procedure. This is because existing rewrite methods select inappropriate number of containers, or ignore the inter-container redundancy information. To address the above problems, we propose an E ffective –Region -Partitioning based rewrite scheme (ERP), which improves the restore performance in backup systems and ensures a high deduplication ratio. The key idea of ERP is to effectively narrow the selection range and choose a flexible number of containers by investigating the inter-container redundancy information. To demonstrate the effectiveness of ERP, we conduct several experiments in a deduplication backup system. Compared to the state-of-the-art rewrite schemes, the results show that ERP reduces the rewrite cost by up to 97.71%. Yihui Lu, Chentao Wu, Jie Li 0002, Minyi Guo |
ICPADS | 3 |
| 2022 | Zero-shot Scene Graph Generation with Relational Graph Neural NetworksabstractExisting scene graph generation (SGG) methods are far from practical, primarily due to their poor performance on predicting zero-shot (i.e., unseen) subject-predicate-object triples. We observe that these SGG methods treat images along with the triples in them independently and thus fail to consider the complex and hidden information that is inherently implicit in the triples of other images. To this effect, our paper proposes a novel encoder-decoder SGG framework to leverage the semantic correlations between the triples of different images into the prediction of a zero-shot triple. Specifically, the encoder aggregates the triples in each image of training set into a large knowledge graph and learns the entity embeddings that capture the features of their neighborhoods with a relational graph neural network. The neighborhood-aware embeddings are then fed into the vision-based decoder to predict the predicates in images. Extensive experiments on the popular benchmark Visual Genome demonstrate that our proposed method outperforms the state-of-the-art methods in popular zero-shot metrics (i.e., zR@N, ngzR@N) for all SGG tasks. Jie Li 0002, Shijing Yuan, Chao Wang 0009, Chentao Wu |
ICPR | 5 |
| 2022 | PRM: An Efficient Partial Recovery Method to Accelerate Training Data Reconstruction for Distributed Deep Learning Applications in Cloud Storage SystemsabstractDistributed deep learning is a typical machine learning method running in distributed environment such as cloud computing systems. The corresponding training, validation and test datasets are very large in general (e.g., several TBs), which need to be stored across multiple data nodes. Due to the high disk failure ratio in cloud storage systems, one of the critical issues for distributed deep learning is how to efficiently tolerate disk failures in the training procedures. These failures can lead to a large amount of data loss, which decreases the training accuracy and slows down the training process. Although several recovery methods are proposed to accelerate the data reconstruction, the related overhead is extremely high, such as high CPU/GPU utilization, a large number of I/Os, etc.To address the above problems, we propose a novel Partial-Recovery Method (called PRM) , which is an adaptive recovery method to accelerate data reconstruction for distributed deep learning applications in cloud storage systems. The key idea of PRM is combining the advantages of erasure coding’s ability to obtain global information on the data distribution with the AI’s ability to recover partial lost data, which can sharply reduce the overhead with acceptable training accuracy. To demonstrate the effectiveness of the PRM approach, we conduct several experiments. The results show that, compared to the state-of-the-art full or approximate recovery methods, PRM decreases the average network transmission time overhead by up to 64.50%, and reduces the recovery time by up to 55.90%, respectively. Piao Hu, Yunfei Gu, Ranhao Jia, Chentao Wu, Minyi Guo, Jie Li 0002 |
IWQoS | 4 |
| 2022 | XHR-Code: An Efficient Wide Stripe Erasure Code to Reduce Cross-Rack Overhead in Cloud Storage SystemsabstractNowadays wide stripe erasure codes (ECs) become popular as they can achieve low monetary cost and provide high reliability for cold data. Generally, wide stripe erasure codes can be generated by extending traditional erasure codes with a large stripe size, or designing new codes. However, although wide stripe erasure codes can decrease the storage cost significantly, the construction of lost data is extraordinary slow, which stems primarily from high cross-rack overhead. It is because a large number of racks participate in the construction of the lost data, which results in high cross-rack traffic. To address the above problems, we propose a novel erasure code called XOR-Hitchhiker-RS (XHR) code, to decrease the cross-rack overhead and still maintain low storage cost. The key idea of XHR is that it utilizes a triple dimensional framework to place more chunks within racks and reduce global repair triggers. To demonstrate the effectiveness of XHR-Code, we provide mathematical analysis and conduct comprehensive experiments. The results show that, compared to the state-of-the-art solutions such as ECWide under various failure conditions, XHR can effectively reduce cross-rack repair traffic and the repair time by up to 36.50%. Guofeng Yang, Huangzhen Xue, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo, Yuanyuan Dong 0002 |
SRDS | 4 |
| 2022 | Measuring Similarity Between Any Pair of Passengers Using Smart Card Usage DataabstractRecent years have witnessed considerable progress in the application of Internet of Things (IoT) technology in smart transportation systems. The wider presence of Wi-Fi networks in subway gates allows passengers to use the quick response (QR) code of mobile phone applications for entrance. The network established by gates has become a medium which connects stations and passengers. However, in addition to directly monitoring the passenger flow, the potential application of the smart card usage data collected by the gates remains an open topic. Although there are several clustering-based works devoted to revealing passengers’ travel behavior patterns, research on the social attributes of subway passengers is very limited. To fill the gap, this article proposes a novel method to mine similarity information of passengers by leveraging passengers’ communication behaviors hidden in subway card usage data. Passengers are first organized as a graph, which not only reflects the interactions between them but also incorporates the context information of subway stations. Then, the node embedding is used to encode the information contained in the graph and with the use of cosine similarity, the similarity between two passengers is measured. Extensive experiments on two real-world location-based social network data sets and extended experiments on a Shanghai subway data set are conducted. The results show that the proposed method can effectively improve the accuracy of similarity measurement and provide social features that are distinguishable from travel behavior patterns. Jie Li 0002, Chentao Wu, Jinsong Wu 0001, Mahmoud Daneshmand |
IEEE Internet Things J. | 3 |
| 2022 | JORA: Blockchain-based efficient joint computing offloading and resource allocation for edge video streaming systems
Shijing Yuan, Jie Li 0002, Chentao Wu |
J. Syst. Archit. | 3 |
| 2021 | Lazy-WL: A Wear-aware Load Balanced Data Redistribution Method for Efficient SSD Array ScalingabstractNowadays, Solid State Drive (SSD) arrays have been widely used in commercial big data centers and high-performance storage services. Meanwhile, in the era of explosive data growth, data centers need to implement the array scaling schemes to meet the increasing storage capacity requirements. The existing state-of-the-art scaling methods, such as Round-Robin (RR) and FastScale, aim at ensuring a uniform data redistribution. However, most of them are designed for Hard Disk Drive (HDD) arrays, ignoring lifetime difference among extended and former-used disks, which leads to several additional penalties in SSD arrays. Furthermore, due to the sudden interdisk lifetime disparity, the extended SSD disks trigger frequently wear-leveling operations for controlling the wearing balance into the predefined threshold. These reactions result in inefficient scaling and I/O performance degradation. To address the above problem, we propose a Lazy W ear-L eveling (Lazy-WL) mechanism to reduce the conventional wear-leveling overhead during the scaling process. Its core idea is to reduce the unnecessary intensive wear-leveling migration significantly, via narrowing the difference of program/erase (P/E) cycles among new-added and former deployed disks smoothly and gradually. To demonstrate the effectiveness of this approach, we conduct several simulation via Disksim and real implementation via a Hadoop cluster. The evaluation results show that, compared to the typical inter and intra disk wear leveling methods, Lazy-WL could lower the triggered wear-leveling operations by up to 92.9% and achieve a maximal 85.2% response time reduction, which suggests that Lazy-WL performs a balanced I/O distribution, and maintains high performance of SSD array with high scaling efficiency. Hanchen Guo, Zhehan Lin, Yunfei Gu, Chentao Wu, Li Jiang 0002, Jie Li 0002, Guangtao Xue, Minyi Guo |
CLUSTER | 4 |
| 2021 | Sharding for Blockchain based Mobile Edge Computing System: A Deep Reinforcement Learning ApproachabstractWith the growth of data scale in the mobile edge computing (MEC) network, data security of the MEC network has become a burning concern. The application of blockchain technology in MEC enhances data security and privacy protection. However, throughput becomes the bottleneck of the blockchain-enabled MEC system. Hence, this paper proposes a novel hierarchical and partitioned blockchain framework to improve scalability while guaranteeing the security of partitions. Next, we model the joint optimization of throughput and security as a Markov decision process (MDP). After that, we adopt deep reinforcement learning (DRL) based algorithms to obtain the number of partitions, the size of micro blocks and the large block generation interval. Finally, we analyze the security and throughput performance of proposed schemes. Simulation results demonstrate that proposed schemes can improve throughput while ensuring the security of partitions. Shijing Yuan, Jie Li 0002, Jinghao Liang, Yuxuan Zhu 0003, Chentao Wu |
GLOBECOM | 7 |
| 2021 | Dense Attention Module for Accurate Pulmonary Nodule DetectionabstractLung cancer has been the leading death cause in modern society. Early detection of pulmonary nodules can significantly improve the survival rate of lung cancer. In this paper, we propose a novel pulmonary nodule detection framework and a novel 3D dense attention module (DAM) which can efficiently exploit the abundant 3D spatial features. The attention module, which integrates the improved dense block and the conv attention block, focuses on three dimensions, plane attention, depth attention, and channel attention. And the whole framework consists of two phases: Nodule Candidate Generation (NCG) and False Positive Reduction (FPR). In NCG phase, we construct a detection network based on DAM. Due to the wide distribution of the nodule diameters, we propose a 3D Feature Pyramid Network (3DFPN) to better handle the scale-varying problem. In FPR phase, we design a 3D DCNN to erase the false positives. Sliding-window based data augment methods are adopted to deal with the unbalance problem of the data. Comprehensive experiments show that our scheme outperforms the existing methods. Jiannan Liu, Jie Li 0002, Fanyong Xue, Chentao Wu |
ICASSP | 4 |
| 2021 | A Novel All-In-One Grid Network for Video Frame InterpolationabstractFlow-based approaches for video frame interpolation typically consist of multiple networks that are responsible for feature extraction, optical flow estimation, and image synthesis, respectively. However, they are usually computationally expensive, and can hardly be employed in devices with limited computing resources. In this work, we propose an All-in-one Grid Frame Interpolation Network (AGFIN) to address this problem. AGFIN is a light-weight network with multiple rows and columns. In each row, we estimate the contextual features and optical flows, then the image synthesis module reconstructs the results from the warped frames and features. Each row serves as a coarser or finer auxiliary for the nearest row. In contrast to using multiple networks, our model integrates feature extraction, optical flow estimation, and image synthesis into a compact network. The experimental results show that our approach has better or comparable performance comparing to representative state-of-the-art approaches with less computational cost. Fanyong Xue, Jie Li 0002, Chentao Wu |
ICIP | 3 |
| 2021 | BWIN: A Bilateral Warping Method for Video Frame InterpolationabstractFlow-based video frame interpolation approaches typically adopt forward or backward warping to approximate the intermediate frames. And a synthesis network is used to refine the interpolation results. Optical flows indicate motion between two input frames, but both forward and backward warping only utilize the first frame. In this work, we propose bilateral warping to make full use of optical flows. Specifically, the proposed bilateral warping yields intermediate candidates from not only the first frame but also the second frame. Our model first applies bilateral warping on the input frames and contextual features. Then, we add skip connections from the input frames and contextual features to the synthesis network. Finally, the synthesis network generates the interpolation results by integrating the original and warped representations. The experimental results on a wide variety of datasets demonstrate the superiority of the proposed approach over the state-of-the-art video frame interpolation methods. Fanyong Xue, Jie Li 0002, Jiannan Liu, Chentao Wu |
ICME | 4 |
| 2021 | Spring Buddy: A Self-Adaptive Elastic Memory Management Scheme for Efficient Concurrent Allocation/Deallocation in Cloud Computing SystemsabstractWithin the cloud computing scenario, each server usually carries multiple service processes, which intensifies the concurrency pressure of the system. As a result, the process of memory management during page allocation and deallocation becomes a significant bottleneck. Although several methods such as Buddy System and Inverse Buddy System (iBuddy) have been proposed to improve the performance of memory management, they cannot adapt to the highly concurrent environment of cloud computing, because they either force the memory allocation/deallocation requests to be serialized or bring extra fragmentation. To address the above problem, we propose Spring Buddy, which improves the concurrency of both memory allocation and deallocation and avoids unnecessary fragmentation. It can detect the changes of system- and process-level memory request patterns and dynamically adjust the organization of page frames. Inventively, Spring Buddy uses the spring core layer to provide both concurrent response and resource aggregation capability which is adapted to the system's concurrency pressure, and also uses the spring lazy layer to further mitigate the system resource contention through process behavior prediction. To demonstrate the effectiveness of Spring Buddy, we implement it in the Linux kernel. The results demonstrate that Spring Buddy can reduce memory allocation latency by 71.47 % and deallocation latency by 93.20% on average compared to the existing methods. Yihui Lu, Chentao Wu, Jia Wang 0009, Xiaoming Gao, Jie Li 0002, Minyi Guo |
ICPADS | 3 |
| 2021 | A Graph-Assisted Out-of-Place Update Scheme for Erasure Coded Storage SystemsabstractErasure Codes (ECs) have widely been used in distributed storage systems to ensure data availability because of its low storage cost and high reliability. However, the update operations in erasure coded storage systems can bring extremely high I/O latency and load imbalance due to the complexity of relationships between data and parity blocks. Although several methods such as Parity Logging (PL) and Log-Structured Array (LSA) have been proposed to improve the performance of updates, they either bring extra I/O operations or decrease the performance of file access. Haiwei Deng, Ranhao Jia, Chentao Wu |
ICPP | 3 |
| 2021 | Rack-Scaling: An efficient rack-based redistribution method to accelerate the scaling of cloud disk arraysabstractIn cloud storage systems, disk arrays are widely used because of their high reliability and low monetary cost. Due to the burst of I/O in sprinting computing scenarios (i.e. online retailer services on Black Friday or Cyber Monday), large scale cloud storage systems such as AWS S3 and GFS need to afford 10XI/O workloads. Therefore, rack level scaling for cloud disk arrays becomes urgent for sprinting services. Although several existing methods, such as Round-Robin(RR) and Scale-RS, are proposed to accelerate the scaling processes, the efficiencies of these approaches are limited. It is because that the cross-rack data migrations are ill-considered in their designs. To address the above problem, in this paper, we propose Rack-Scaling, a novel data redistribution method to accelerate rack level scaling process in cloud storage systems. The basic idea of Rack-Scaling is migrating appropriate data blocks within and among racks to achieve a uniform data distribution while minimizing the cross-rack migration, which costs more than intra-rack migration. We conduct simulations via Disksim and we also implement Rack-Scaling on Hadoop to demonstrate the effectiveness of Rack-Scaling. The results show that, compared to typical methods such as Round-Robin (RR), Semi-RR, Scale-RS and BDR, Rack-Scaling reduces the number of I/O operations and the data amount of cross-rack transmission by up to 90.4% and 99.9%, respectively, and speeds up the scaling by up to 8.77X. Zhehan Lin, Hanchen Guo, Chentao Wu, Jie Li 0002, Guangtao Xue, Minyi Guo |
IPDPS | 3 |
| 2021 | EC-Scheduler: A Load-Balanced Scheduler to Accelerate the Straggler Recovery for Erasure Coded Storage SystemsabstractErasure codes (EC) have become a typical technology for distributed storage systems in place of data replication, providing similar data availability but lower storage cost. However, a great number of data computations and migrations during the EC recovery process bring high I/O and network latency penalties. Although several EC recovery methods have been designed to compromise the recovery penalty with high parallelism, the performance of these schemes was usually bounded by the straggler problems due to the various (I/O) performance among different nodes in the storage system. Moreover, the variation of the access popularity from the upper layer application causes the dynamic load fluctuation and asymmetry upon different nodes, which makes the scheduling more difficult during the recovery. To address the above problem, we propose a dynamic load-balanced scheduling algorithm for straggler recovery called EC-Scheduler. EC-Scheduler adjusts the recovery schedule dynamically with the awareness of continuous load fluctuation on the nodes, guaranteeing high parallelism and load balance ability simultaneously. To demonstrate the effectiveness of EC-Scheduler, we conduct several experiments in a cluster. The results show that, compared to typical recovery schemes such as Fast-PR and EC-Store, EC-Scheduler could achieve a 1.3X speed-up in the recovery process and 10X improvement in recovery load imbalance factor. Xinzhe Cao, Yunfei Gu, Chentao Wu, Jie Li 0002, Guangtao Xue, Minyi Guo, Yuanyuan Dong 0002 |
IWQoS | 4 |
| 2020 | FAGR: An Efficient File-aware Graph Recovery Scheme for Erasure Coded Cloud Storage SystemsabstractWith the explosive growth of data in cloud storage systems, Erasure Codes (ECs) have become a typical data redundancy technology because of its low storage cost and high reliability. However, due to a large amount of complex computations and transmissions among massive data and parities, the recovery of lost data in erasure coded storage systems incurs high I/O latency. Although several fast recovery approaches devote to mitigating the recovery time from the application level or device level, the performance of file level recovery is still restricted. It is because a part of the complicated relationships among data, parity and files are ignored in the design of recovery process. To address the above problems, we propose a novel File-aware Graph Recovery (FAGR) scheme, to improve the file level recovery performance during the reconstruction process. The key idea of FAGR is establishing a graph with the mappings among files, blocks, stripes, parities, nodes and the access frequencies of files, and guides the recovery process from file point of view. A corresponding model is established to analyze the cost efficiency of recovery process, which guarantees that FAGR reconstructs the popular files in advance to accelerate the recovery. To demonstrate the effectiveness of FAGR, we conduct several numerical analysis and experiments in clusters. The results show that, compared to typical fast recovery methods, FAGR reduces the average response time of files by up to 81.63 % and improves the throughput by up to 4.44 ×. Heming Zeng, Chi Zhang 0005, Chentao Wu, Jie Li 0002, Guangtao Xue, Minyi Guo |
ICCD | 3 |
| 2020 | DCVP: Distributed Collaborative Video Stream Processing in Edge ComputingabstractIn edge computing, computation offloading of video stream tasks and collaboration processing among edge nodes is a huge challenge. The previous research mainly focuses on the selection of computing modes and resource allocation, but taking no joint consideration of computation offloading and collaborative processing of edge node groups. In order to jointly tackle these issues in edge computing, we propose an innovative distributed collaborative video stream processing framework for edge computing(DCVP), where the video tasks are assigned to mobile edge computing (MEC) nodes or edge groups based on the offloading decision. First, we design a method for the group formation, which matches video subtasks to appropriate edge groups. In addition, we present two offloading modes for video streaming tasks, e.g., offloading to MEC nodes or edge groups, to handle computationally intensive video tasks. Furthermore, we formulate the joint optimization problem for offloading decision and collaborative processing of video subtasks into a distributed optimization problem. Finally, we employ an alternating direction method of multipliers (ADMM)-based algorithm to solve the problem. Simulation results under multiple parameters show the proposed schemes outperform other typical schemes. Shijing Yuan, Jie Li 0002, Chentao Wu, Yusheng Ji, Yongbing Zhang 0001 |
ICPADS | 3 |
| 2020 | Small Object Detection by Generative and Discriminative LearningabstractWith the development of deep convolutional neural networks (CNNs), the object detection accuracy has been greatly improved. But the performance of small object detection is still far from satisfactory, mainly because small objects are so tiny that the information contained in the feature map is limited. Existing methods focus on improving classification accuracy but still suffer from the limitation of bounding box prediction. To solve this issue, we propose a detection framework by generative and discriminative learning. First, a reconstruction generator network is designed to reconstruct the mapping from low frequency to high frequency for anchor box prediction. Then, a detector module extracts the regions of interest (ROIs) from generated results and implements a RoI-Head to predict object category and refine bounding box. In order to guide the reconstructed image related to the corresponding one, a discriminator module is adopted to tell from the generated result and the original image. Extensive evaluations on the challenging MS-COCO dataset demonstrate that our model outperforms most state-of-the-art models in detecting small objects, especially the reconstruction module improves the average precision for small object (APs) by 7.7%. Jie Li 0002, Chentao Wu, Weijia Jia 0001 |
ICPR | 3 |
| 2020 | EC-Fusion: An Efficient Hybrid Erasure Coding Framework to Improve Both Application and Recovery Performance in Cloud Storage SystemsabstractNowadays erasure coding is one of the most significant techniques in cloud storage systems, which provides both quick parallel I/O processing and high capabilities of fault tolerance on massive data accesses. In these systems, triple disk failure tolerant arrays (3DFTs) is a typical configuration, which is supported by several classic erasure codes like Reed-Solomon (RS) codes, Local Reconstruction Codes (LRC), Minimum Storage Regeneration (MSR) codes, etc. For an online recovery process, the foreground application workloads and the background recovery workloads are handled simultaneously, which requires a comprehensive understanding on both two types of workload characteristics. Although several techniques have been proposed to accelerate the I/O requests of online recovery processes, they are typically unilateral due to the fact that the above two workloads are not combined together to achieve high cost-effective performance.To address this problem, we propose Erasure Codes Fusion (EC-Fusion), an efficient hybrid erasure coding framework in cloud storage systems. EC-Fusion is a combination of RS and MSR codes, which dynamically selects the appropriate code based on its properties. On one hand, for write-intensive application workloads or low risk on data loss in recovery workloads, EC-Fusion uses RS code to decrease the computational overhead and storage cost concurrently. On the other hand, for read-intensive or frequent reconstruction in workloads, MSR code is a proper choice. Therefore, a better overall application and recovery performance can be achieved in a cost-effective fashion. To demonstrate the effectiveness of EC-Fusion, several experiments are conducted in hadoop systems. The results show that, compared with the traditional hybrid erasure coding techniques, EC-Fusion accelerates the response time for application by up to 1.77×, and reduces the reconstruction time by up to 69.10%. Han Qiu 0003, Chentao Wu, Jie Li 0002, Minyi Guo, Tong Liu 0030, Xubin He, Yuanyuan Dong 0002 |
IPDPS | 2 |
| 2020 | AZ-Recovery: An Efficient Crossing-AZ Recovery Scheme for Erasure Coded Cloud Storage SystemsabstractAs massive data in modern cloud storage systems grow dramatically, it is a common method to partition and store data in multiple Availability Zones (AZs). Multiple AZs not only provide high reliability, but also reduce the network latency. Erasure Codes (ECs) are widely used in multiple AZs to provide high reliability at low storage cost. However, the recovery cost of EC is extremely high in multiple AZs' environment, which is mainly because a normal EC needs to reconstruct the lost data via transferring the data/parities across AZs. Although existing fast recovery approaches can save the I/O cost or network bandwidth in an effective manner, they are not suitable for multiple AZs. The reasons include low flexibility on various complex network scenarios, less consideration on crossing-AZ bandwidth, low capabilities on multiple disk/node failures, etc. To address the above problem, in this paper, we propose a crossing $\underline{\mathrm{A}}$vailability Zone Recovery (AZ-Recovery) method to efficiently improve the recovery performance for multiple AZs. AZ-Recovery investigates the complex homogeneous/heterogeneous network topologies, and finds an optimal data transmission path. Using this method, AZ-Recovery can significantly reduce the recovery cost and save the crossing AZ bandwidth in various failure scenarios. To demonstrate the effectiveness of AZ-Recovery, we evaluate various erasure codes via mathematical analysis and simulations in Network Simulator-3. The results show that, compared to the traditional erasure coding methods, AZ-Recovery saves the recovery bandwidth by up to 77.47%. Chentao Wu, Zongxin Ye, Xubin He, Jie Li 0002, Minyi Guo, Guangtao Xue, Yuanyuan Dong 0002 |
SRDS | 2 |
| 2020 | AIR: an approximate intelligent redistribution approach to accelerate RAID scaling
Zhehan Lin, Hanchen Guo, Chentao Wu |
CCF Trans. High Perform. Comput. | 3 |
| 2020 | A pure hardware-driven scheduler for enhancing bank-level parallelism in a persistent memory controller
Dongliang Xue, Linpeng Huang, Chentao Wu |
Future Gener. Comput. Syst. | 3 |
| 2019 | An Efficient Massive Log Discriminative Algorithm for Anomaly Detection in CloudabstractLog anomaly detection is a critical step towards building a secure and trustworthy cloud system. As more corporations turn to cloud system to store and process their most valuable data, the risk of a potential breach of those systems increases exponentially. However, conventional top-n log candidates anomaly detection methods, such as Deeplog and N-gram, often suffer from the limited scope of the top-n list, which rules out many potentially suitable candidates. In this paper, we propose Discounted Cumulative Gain (DCG) discriminative algorithm that ranks all the log candidates and calculates the dcg score to determine the number of log candidates. To demonstrate the effectiveness of our algorithm, we conduct comprehensive experiments under different log workloads. Experimental evaluations show that DCG has outperformed Deeplog and N-gram methods in cloud systems, and improved the F-score of Deeplog and N-gram by up to 3.8% and 11.6% respectively. Jie Li 0002, Chentao Wu |
GLOBECOM | 3 |
| 2019 | DeepDDoS: Online DDoS Attack DetectionabstractHighly efficient and dependable large-scale DDoS attack detection scheme is critical for network anomaly detection. Typical machine learning algorithms such as Decision Tree and Adaboost work well on flow level analysis but cannot perform fine- grained detection of packet levels. Since these algorithms require more packets information for detection, resulting in higher detection delay and relatively lower accuracy. To address the problem, we propose DeepDDoS which is a deep learning method focusing on both period- wise and packetwise attack detection. First, the network packets are modeled in time dimension to discover the potential abnormal time period. Second, the network packets are grouped by 5 tuples (flow), the packets inside the group are sorted according to their arrival time. Then the data packet level sequence modeling is performed in each group. Comprehensive performance evaluation shows that the detection accuracy of DeepDDoS reach 99%. Furthermore, only 5 consecutive packets are needed for packet-wise detection, greatly reducing detection delay and computational overhead. Comparative experiments show that DeepDDoS outperforms existing typical attack detection methods. Zhenping Shi, Jie Li 0002, Chentao Wu |
GLOBECOM | 3 |
| 2019 | Approximate Code: A Cost-Effective Erasure Coding Framework for Tiered Video Storage in Cloud SystemsabstractNowadays massive video data are stored in cloud storage systems, which are generated by various applications such as autonomous driving, news media, security monitoring, etc. Meanwhile, erasure coding is a popular technique in cloud storage to provide both high reliability and low monetary cost, where triple disk failure tolerant arrays (3DFTs) is a typical choice. Therefore, how to minimize the storage cost of video data in 3DFTs is a challenge for cloud storage systems. Although there are several solutions like approximate storage technique, they cannot guarantee low storage cost and high data reliability concurrently. Huayi Jin, Chentao Wu, Jie Li 0002, Minyi Guo |
ICPP | 2 |
| 2019 | Optimizing the Parity Check Matrix for Efficient Decoding of RS-Based Cloud Storage SystemsabstractIn large scale distributed systems such as cloud storage systems, erasure coding is a fundamental technique to provide high reliability at low monetary cost. Compared with the traditional disk arrays, cloud storage systems use an erasure coding scheme with both flexible fault tolerance and high scalability. Thus, Reed-Solomon (RS) Codes or RS-based codes are popular choices for cloud storage systems. However, the decoding performance for RS-based codes is not as good as XOR-based codes, which are optimized via investigating the relationships among different parity chains or reducing the computational complexity of matrix multiplications. Therefore, exploring an efficient decoding method is highly desired. To address the above problem, in this paper, we propose an Advanced Parity-Check Matrix (APCM) based approach, which is extended from the original Parity-Check Matrix based (PCM) approach. Instead of improving the decoding performance of XOR-based codes in PCM, APCM focuses on optimizing the decoding efficiency for RS-based codes. Furthermore, APCM avoids the matrix inversion computations and reduces the computational complexity of the decoding process. To demonstrate the effectiveness of the APCM, we conduct intensive experiments by using both RS-based and XOR-based codes under cloud storage environment. The results show that, compared to typical decoding methods, APCM improves the decoding speed by up to 32.31% in the Alibaba cloud storage system. Junqing Gu, Chentao Wu, Han Qiu 0003, Jie Li 0002, Minyi Guo, Xubin He, Yuanyuan Dong 0002 |
IPDPS | 2 |
| 2019 | AZ-Code: An Efficient Availability Zone Level Erasure Code to Provide High Fault Tolerance in Cloud Storage SystemsabstractAs data in modern cloud storage system grows dramatically, it's a common method to partition data and store them in different Availability Zones (AZs). Multiple AZs not only provide high fault tolerance (e.g., rack level tolerance or disaster tolerance), but also reduce the network latency. Replication and Erasure Codes (EC) are typical data redundancy methods to provide high reliability for storage systems. Compared with the replication approach, erasure codes can achieve much lower monetary cost with the same fault-tolerance capability. However, the recovery cost of EC is extremely high in multiple AZ environment, especially because of its high bandwidth consumption in data centers. LRC is a widely used EC to reduce the recovery cost, but the storage efficiency is sacrificed. MSR code is designed to decrease the recovery cost with high storage efficiency, but its computation is too complex. To address this problem, in this paper, we propose an erasure code for multiple availability zones (called AZ-Code), which is a hybrid code by taking advantages of both MSR code and LRC codes. AZ-Code utilizes a specific MSR code as the local parity layout, and a typical RS code is used to generate the global parities. In this way, AZ-Code can keep low recovery cost with high reliability. To demonstrate the effectiveness of AZ-Code, we evaluate various erasure codes via mathematical analysis and experiments in Hadoop systems. The results show that, compared to the traditional erasure coding methods, AZ-Code saves the recovery bandwidth by up to 78.24%. Chentao Wu, Junqing Gu, Han Qiu 0003, Jie Li 0002, Minyi Guo, Xubin He, Yuanyuan Dong 0002 |
MSST | 2 |
| 2019 | Exploring Transfer Learning to Reduce Training Overhead of HPC Data in Machine LearningabstractNowadays, scientific simulations on high-performance computing (HPC) systems can generate large amounts of data (in the scale of terabytes or petabytes) per run. When this huge amount of HPC data is processed by machine learning applications, the training overhead will be significant. Typically, the training process for a neural network can take several hours to complete, if not longer. When machine learning is applied to HPC scientific data, the training time can take several days or even weeks. Transfer learning, an optimization usually used to save training time or achieve better performance, has potential for reducing this large training overhead. In this paper, we apply transfer learning to a machine learning HPC application. We find that transfer learning can reduce training time without, in most cases, significantly increasing the error. This indicates transfer learning can be very useful for working with HPC datasets in machine learning applications. Tong Liu 0030, Shakeel Alibhai, Jinzhen Wang, Qing Liu 0002, Xubin He, Chentao Wu |
NAS | 6 |
| 2019 | PAM: an efficient power-aware multilevel cache policy to reduce energy consumption of storage systems
Xiaodong Meng, Chentao Wu, Minyi Guo, Long Zheng 0001 |
Frontiers Comput. Sci. | 2 |
| 2019 | Dapper: An Adaptive Manager for Large-Capacity Persistent MemoryabstractIn-memory computing has inspired researchers to consider integrating large-capacity persistent memory (PM) into the main memory subsystem. However, several challenges still remain for providing an integration approach for DRAM-comparable PM on existing enterprise servers. Current commercial servers tend to feature multiple sockets with shared-memory NUMA organizations. Simply constructing a hybrid main memory architecture for these NUMA organizations requires considerable modifications of the system software. Another significant problem in these designs is the high latency of accessing PM on a remote socket, which results in performance degradation. To address these problems, we integrated PM as a memory-based model and as a storage-based model simultaneously on one commercial server, which offers a short-cut approach for enterprises to build commercial NUMA machines with large-capacity PM. In the memory-based model, rather than focusing on the persistence attribute, we propose an architecture that benefits managing the integrated PM and DRAM space in a unified manner and that facilitates bypassing vast modifications to the system software. We also present an adaptive mechanism that can automatically introduce a moderate amount of PM into the local socket to hinder access of a remote socket by the degree of memory pressure. In the storage-based model, under the condition of taking full advantage of the PM's persistence, we abstract a PM volume device and overcome the torn sector problem. To demonstrate the effectiveness of the proposed scheme, we design and implement Dapper, an adaptive persistent memory manager prototype. The experimental results show that, compared to typical memory management approaches, Dapper achieves performance improvements of 13.1 percent to 34.0 percent on average on Graph500 BFS_SSSP benchmarks and SPEC CPU2006 floating point workloads, respectively. Moreover, when deploying F2FS on our PM volume, we find that Dapper outperforms existing methods by 5.8 percent on tar and by 11.9 percent on untar. Dongliang Xue, Linpeng Huang, Chao Li 0009, Chentao Wu |
IEEE Trans. Computers | 4 |
| 2018 | WarmCache: A Comprehensive Distributed Storage System Combining Replication, Erasure Codes and Buffer Cache
Brian A. Ignacio, Chentao Wu, Jie Li 0002 |
GPC | 2 |
| 2018 | Adaptive Memory Fusion: Towards Transparent, Agile Integration of Persistent MemoryabstractThe great promise of in-memory computing inspires engineers to scale their main memory subsystems in a timely and efficient manner. Offering greatly expanded capacity at near-DRAM speed, today's new-generation persistent memory (PM) module is no doubt an ideal candidate for system upgrade. However, integrating DRAM-comparable PMs in current enterprise systems faces big barriers in terms of huge system modifications for software compatibility and complex runtime support. In addition, the very large PM capacity unavoidably results in massive metadata, which introduces significant performance and energy overhead. The inefficiency issue becomes even acute when the memory system reaches its capacity limit or the application requires large memory space allocation. In this paper we propose adaptive memory fusion (AMF), a novel PM integration scheme that jointly solves the above issues. Rather than struggle to adapt to the persistence property of PM through modifying the full software stack, we focus on exploiting the high capacity feature of emerging PM modules. AMF is designed to be totally transparent to user applications by carefully hiding PM devices and managing the available PM space in a DRAM-like way. To further improve the performance, we devise holistic optimization scheme that allows the system to efficiently utilize system resources. Specifically, AMF is able to adaptively release PM based on memory pressure status, smartly reclaim PM pages, and enable fast space expansion with direct PM pass-through. We implement AMF as a kernel subsystem in Linux. Compared to traditional approaches, AMF could decrease the page faults number of high-resident-set benchmarks by up to 67.8% with an average of 46.1%. Using realistic in-memory database, we show that AMF outperforms existing solutions by 57.7% on SQLite and 21.8% on Redis. Overall, AMF represents a more lightweight design approach and it would greatly encourage rapid and flexible adoption of PM in the near future. Dongliang Xue, Chao Li 0009, Linpeng Huang, Chentao Wu, Tianyou Li |
HPCA | 4 |
| 2018 | Reference-Counter Aware Deduplication in Erasure-Coded Distributed Storage SystemabstractIn modern distributed storage systems, space efficiency and system reliability are two major concerns. As a result, contemporary storage systems often employ data deduplication and erasure coding to reduce the storage overhead and provide fault tolerance, respectively. However, little work has been done to explore the relationship between these two techniques. In this paper, we propose Reference-counter Aware Deduplication (RAD), which employs the features of deduplication into erasure coding to improve garbage collection performance when deletion occurs. RAD wisely encodes the data according to the reference counter, which is provided by the deduplication level and thus reduces the encoding overhead when garbage collection is conducted. Further, since the reference counter also represents the reliability levels of the data chunks, we additionally made some effort to explore the trade-offs between storage overhead and reliability level among different erasure codes. The experiment results show that RAD can effectively improve the GC performance by up to 24.8% and the reliability analysis shows that, with certain data features, RAD can provide both better reliability and better storage efficiency compared to the traditional Round- Robin placement. Tong Liu 0030, Xubin He, Shakeel Alibhai, Chentao Wu |
NAS | 4 |
| 2018 | Toward multi-programmed workloads with different memory footprints: a self-adaptive last level cache scheduling scheme
Minyi Guo, Chentao Wu |
Sci. China Inf. Sci. | 3 |
| 2018 | HSCS: a hybrid shared cache scheduling scheme for multiprogrammed workloads
Chentao Wu, Dingyu Yang, Xiaodong Meng, Liting Xu, Minyi Guo |
Frontiers Comput. Sci. | 2 |
| 2017 | Loc-K: A Spatial Locality-Based Memory Deduplication Scheme with Prediction on K-Step LocationsabstractMemory deduplication is a technique to eliminate redundant data, save memory space and improve the performance of the whole system. There are several effective deduplication algorithms, which identify replicated data via comparing the content of different pages. However, although a few literatures utilize spatial locality to improve the efficiency of memory deduplication [1][2], they still have several limitations, such as low ratio of continuous distribution and high failure rate of prediction under unstable environments. To address these problems, in this paper, we design a new memory deduplication algorithm called “Loc-K”. On one hand, it utilizes logical addresses of different pages to ensure better continuity, which can gain better spatial locality. On the other hand, Loc-K predicts K potential duplication locations as the targets for page scanning, which improves the prediction hit ratio. Furthermore, Loc-K merges the duplicated pages directly to avoid regular searching routines. To demonstrate the effectiveness of our algorithm, we conduct several experimentations via implementation in Linux Kernel. The results show that, compared to the state-of-the-art memory deduplication algorithms, Loc-K increases the predictable opportunity by up to 97.8%, increases the prediction hit ratio by up to 96.5%, and reduces the duplication identification time by at least 34.3% respectively. Shuaijie Jia, Chentao Wu, Jie Li 0002 |
ICPADS | 2 |
| 2017 | Favorable Block First: A Comprehensive Cache Scheme to Accelerate Partial Stripe Recovery of Triple Disk Failure Tolerant ArraysabstractWith the development of cloud computing, disk arrays tolerating triple disk failures (3DFTs) are receiving more attention nowadays because they can provide high data reliability with low monetary cost. However, a challenging issue in these arrays is how to efficiently reconstruct the lost data, especially for partial stripe errors (e.g., sector and chunk errors). It is one of the most significant scenarios in practice. However, existing cache strategies are not efficient for partial stripe reconstruction in 3DFTs, which is because the complex relationships among data and parities are usually ignored during the recovery process. To address this problem, in this paper, we proposed a comprehensive cache policy called Favorable Block First (FBF), which can speed up the partial stripe reconstruction of 3DFTs. FBF investigates the relationships among parity chains via allocating various priorities of shared chunks. Thus in the recovery process, by giving higher priorities to the chunks which are shared by more parities chains, FBF can dynamically hold the significant data in buffer cache for partial stripe reconstruction. Obviously, it increases the cache hit ratio and reduces the reconstruction time. To demonstrate the effectiveness of FBF, we conduct several simulations via Disksim. The results show that, compared to typical recovery schemes by combining with classic cache policies (e.g., LRU, LFU and ARC), FBF improves hit ratio by up to 2.47 times and accelerates the reconstruction process by 14.90%, respectively. Luyu Li, Houxiang Ji, Chentao Wu, Jie Li 0002, Minyi Guo |
ICPP | 3 |
| 2017 | A Hint Frequency Based Approach to Enhancing the I/O Performance of Multilevel Cache Storage Systems
Xiaodong Meng, Chentao Wu, Minyi Guo, Jie Li 0002, Xiaoyao Liang, Bin Yao 0002, Long Zheng 0001 |
J. Comput. Sci. Technol. | 2 |
| 2016 | BDR: A Balanced Data Redistribution scheme to accelerate the scaling process of XOR-based Triple Disk Failure Tolerant arraysabstractIn large scale data centers, with the increasing amount of user data, Triple Disk Failure Tolerant arrays (3DFTs) gain much popularity due to their high reliability and low monetary cost. With the development of cloud computing, scalability becomes a challenging issue for disk arrays like 3DFTs. Although previous solutions improves the efficiency of RAID scaling, they suffer many problems (high I/O overhead and long migration time) in 3DFTs. It is because that existing approaches have to cost plenty of migration I/Os on balancing the data distribution according to the complex layout of erasure codes. To address this problem, we propose a novel Balanced Data Redistribution scheme (BDR) to accelerate the scaling process, which can be applied on XOR-based 3DFTs. BDR migrates proper data blocks according to a global point of view on a stripe set, which guarantees uniform data distribution and a small number of data movements. To demonstrate the effectiveness of BDR, we conduct several evaluations and simulations. The results show that, compared to typical RAID scaling approaches like Round-Robin (RR), SDM and RS6, BDR reduces the scaling I/Os by up to 77.45%, which speeds up the scaling process of 3DFTs by up to 4.17×, 3.31×, 3.88×, respectively. Yanbing Jiang, Chentao Wu, Jie Li 0002, Minyi Guo |
ICCD | 2 |
| 2016 | Zero-Chunk: An Efficient Cache Algorithm to Accelerate the I/O Processing of Data DeduplicationabstractData deduplication is a technique to eliminate duplicated copies of data. It can save the storage space, reduce the amount of disk I/Os, then improve the system performance. There have been several popular deduplication algorithms such as SISL [30], Extreme Binning [1], Sparse Indexing [14], etc. These schemes use containers to aggregate data chunks for better performance. However, they either suffer from low cache hit ratios or inefficient cache utilization. To address this problem, we design Zero-Chunk, a new cache algorithm that balances the cache hit ratio and memory usage. In our method, we choose chunks whose fingerprints have all-zero remainders as pointers (called zero chunks), and aggregate the following chunks into their corresponding containers. And then, when the access patterns change, our method can eliminate cold data chunks and containers to maintain a low overhead. To demonstrate the effectiveness of Zero-Chunk, we conduct several simulations. The results show that, compared to Sparse Indexing (the most popular implementation method in data deduplication), Zero-Chunk improves the cache hit ratio by up to 5.2%, saves the memory consumption by more than 50.7%, and decreases the total number of I/Os by up to 17.3%, respectively. Hongyuan Gao, Chentao Wu, Jie Li 0002, Minyi Guo |
ICPADS | 2 |
| 2016 | DASM: A Dynamic Adaptive Forward Assembly Area Method to Accelerate Restore Speed for Deduplication-Based Backup Systems
Luyu Li, Chentao Wu, Jie Li 0002 |
NPC | 3 |
| 2015 | BPS: A Balanced Partial Stripe Write Scheme to Improve the Write Performance of RAID-6abstractNowadays RAID is widely used due to its large capacity, high performance and high reliability. With the increasing requirement of reliability in storage systems and fast development of cloud computing, RAID-6, which can tolerate concurrent failures of any two disks, receives more attention than ever. However, the write performance of RAID-6 systems is a bottleneck to serve various applications. In the last two decades, many approaches are proposed to enhance the write performance of RAID-6, but they have several limitations, such as unbalanced I/O distribution and high I/O cost. To address this problem, in this paper, we propose a Balanced Partial Stripe (BPS) write scheme to improve the write performance of RAID-6 systems. The basic idea of BPS is reorganizing the distribution of write data blocks according to a global point of view on modified parities, and flushing these blocks to storage devices at once. Therefore, it can significantly reduce the total number of parity updates and balance the I/O workload. BPS has three main advantages: 1) BPS decreases the number of I/O operations and aggregate the fragmented I/Os, which improves the I/O performance, 2) BPS provides a balanced partial stripe write approach for RAID-6, 3) BPS can be applied with various erasure codes. To demonstrate the effectiveness of our scheme, we conduct simulations on DiskSim to evaluate different partial stripe write approaches. The results show that, compared to typical partial stripe write approaches, BPS reduces the average access time by up to 37.14%, and decreases the number of write operations by up to 26.24%. Congjin Du, Chentao Wu, Jie Li 0002, Minyi Guo, Xubin He |
CLUSTER | 2 |
| 2015 | TIP-Code: A Three Independent Parity Code to Tolerate Triple Disk Failures with Optimal Update ComplextiyabstractWith the rapid expansion of data storages and the increasing risk of data failures, triple Disk Failure Tolerant arrays (3DFTs) become popular and widely used. They achieve high fault tolerance via erasure codes. One class of erasure codes called Maximum Distance Separable (MDS) codes, which aims to offer data protection with minimal storage overhead, is a typical choice to enhance the reliability of storage systems. However, existing 3DFTs based on MDS codes are inefficient in terms of update complexity, which results in poor write performance. In this paper, we present an efficient MDS coding scheme called TIP-code, which is purely based on XOR operations and can tolerate triple disk failures. It uses three independent parities (horizontal, diagonal and anti-diagonal parities), and offers optimal update complexity. To demonstrate the effectiveness of TIP-code, we conduct several quantitative analysis and experiments. The results show that, compared to typical MDS codes for 3DFTs (i.e., Cauchy-RS and STAR codes), TIP-code improves the single write performance by up to 46.6%. Yongzhe Zhang, Chentao Wu, Jie Li 0002, Minyi Guo |
DSN | 2 |
| 2015 | EH-Code: An Extended MDS Code to Improve Single Write Performance of Disk Arrays for Correcting Triple Disk Failures
Yanbing Jiang, Chentao Wu, Jie Li 0002, Minyi Guo |
ICA3PP (1) | 2 |
| 2015 | FDRC: Flow-driven rule caching optimization in software defined networkingabstractWith the sharp growth of cloud services and their possible combinations, the scale of data center network traffic has an inevitable explosive increasing in recent years. Software defined network (SDN) provides a scalable and flexible structure to simplify network traffic management. It has been shown that Ternary Content Addressable Memory (TCAM) management plays an important role on the performance of SDN. However, previous literatures, in point of view on rule placement strategies, are still insufficient to provide high scalability for processing large flow sets with a limited TCAM size. So caching is a brand new method for TCAM management which can provide better performance than rule placement. In this paper, we propose FDRC, an efficient flow-driven rule caching algorithm to optimize the cache replacement in SDN-based networks. Different from the previous packet-driven caching algorithm, FDRC is characterized by trying to deal with the challenges of limited cache size constraint and unpredictable flows. In particular, we design a caching algorithm with low-complexity to achieve high cache hit ratio by prefetching and special replacement strategy for predictable and unpredictable flows, respectively. By conducting extensive simulations, we demonstrate that our proposed caching algorithm significantly outperforms FIFO and least recently used (LRU) algorithms under various network settings. He Li 0001, Song Guo 0001, Chentao Wu, Jie Li 0002 |
ICC | 3 |
| 2015 | Code 5-6: An Efficient MDS Array Coding Scheme to Accelerate Online RAID Level MigrationabstractWith the rapid growth of data storage, the demand for high reliability becomes critical in large data centers where RAID-5 is widely used. However, the disk failure rate increases sharply after some usage, and thus concurrent disk failures are not rare, therefore RAID-5 is insufficient to provide high reliability. A solution is to convert an existing RAID-5 to a RAID-6 (a type of "RAID level migration") to tolerate more concurrent disk failures via erasure codes, but existing approaches involve complex conversion process and high transformation cost. To address these challenges, we propose a novel MDS code, called "Code 5-6", to combine a new dedicated parity column with the original RAID-5 layout. Code 5-6 not only accelerates online conversion from a RAID-5 to a RAID-6, but also demonstrates several optimal properties of MDS codes. Our mathematical analysis shows that, compared to existing MDS codes, Code 5-6 reduces new parities, decreases the total I/O operations, and speeds up the conversion process by up to 80%, 48.5%, and 3.38×, respectively. Chentao Wu, Xubin He, Jie Li 0002, Minyi Guo |
ICPP | 1 |
| 2015 | PCM: A Parity-Check Matrix Based Approach to Improve Decoding Performance of XOR-based Erasure CodesabstractIn large storage systems, erasure codes is a primary technique to provide high reliability with low monetary cost. Among various erasure codes, a major category called XORbased codes uses purely XOR operations to generate redundant data and offer low computational complexity. These codes are conventionally implemented via matrix based method or several specialized non-matrix based methods. However, these approaches are insufficient on decoding performance, which affects the reliability and availability of storage systems. To address the problem, in this paper, we propose a novel Parity-Check Matrix based (PCM) approach, which is a general-purpose method to implement XOR-based codes, and increases the decoding performance by using smaller and sparser matrices. To demonstrate the effectiveness of PCM, we conduct several experiments by using different XOR-based codes. The evaluation results show that, compared to typical matrix based decoding methods, PCM can improve the decoding speed by up to a factor of 1.5× when using EVENODD code (an erasure code for correcting double disk failures), and accelerate the decoding process of STAR code (an erasure code for correcting triple disk failures) by up to a factor of 2.4×. Yongzhe Zhang, Chentao Wu, Jie Li 0002, Minyi Guo |
SRDS | 2 |
| 2014 | An Advanced Data Redistribution Approach to Accelerate the Scale-Down Process of RAID-6
Congjin Du, Chentao Wu, Jie Li 0002 |
ICA3PP (2) | 2 |
| 2014 | HFA: A Hint Frequency-based approach to enhance the I/O performance of multi-level cache storage systemsabstractWith the enormous and increasing user demand, I/O performance is one of the primary considerations to build a data center. Several new technologies in data centers, such as tiered storage [33], prompt the widespread usage of multi-level cache techniques. In these storage systems, the upper level storage typically serves as a cache for the lower level, which forms a distributed multi-level cache system. However, although many excellent multi-level cache algorithms are proposed to improve the I/O performance, they still have potential to be enhanced by investigating the history information of hints [28]. To address this challenge, in this paper, we propose a novel Hint Frequency-based Approach (HFA), to improve the overall multi-level cache performance of storage systems. The main idea of HFA is using hint frequencies (the total number of demotions/promotions by employing demote/promote hints) to efficiently explore the valuable history information of data blocks among multiple levels. HFA can be applied with several popular multi-level cache algorithms, such as Demote, Promote, Hint-K, etc. Simulation results show that, compared to original multi-level cache algorithms such as Demote, Promote and Hint-K, HFA can improve the I/O performance by up to 20% under different I/O workloads. Xiaodong Meng, Chentao Wu, Jie Li 0002, Xiaoyao Liang, Bin Yao 0002, Minyi Guo, Long Zheng 0001 |
ICPADS | 2 |
| 2014 | LSShare: an efficient multiple query optimization system in the cloud
Xing Ge, Bin Yao 0002, Minyi Guo, Changliang Xu, Jingyu Zhou, Chentao Wu, Guangtao Xue |
Distributed Parallel Databases | 6 |
| 2014 | Hint-K: An Efficient Multilevel Cache Using K-Step HintsabstractI/O performance has been critical for large-scale distributed systems. Many approaches, including hint-based multilevel cache, have been proposed to smooth the gap between different levels. These solutions demote or promote cache blocks based on the latest history information, which is insufficient for applications where frequent demote and promote operations occur. In this paper, we propose a novel multilevel buffer cache using K-step hints (Hint-K) to improve the I/O performance of distributed systems. The basic idea is to promote a block from the lower level cache to the higher level(s) or demote a block vice versa based on the block's previous K-step promote or demote operations, which are referred to as K-step hints. If we make an analogy between Hint-K and LRU-K, then LRU-K keeps track of the times of last K references for blocks within a single cache level, while our Hint-K keeps track of the information of the last K movements (either demote or promote) of blocks among different cache levels. We develop our Hint-K algorithms and design a mathematical model that can efficiently describe the activeness of any block in any cache level. Simulation results show that Hint-K achieves better performance compared to the existing popular multilevel cache schemes such as PROMOTE, DEMOTE, and MQ under different I/O workloads. Chentao Wu, Xubin He, Qiang Cao 0001, Changsheng Xie 0001, Shenggang Wan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | A Flexible Framework to Enhance RAID-6 Scalability via Exploiting the Similarities among MDS CodesabstractWith increasing demand in high performance and reliability, RAID systems especially RAID-6 are widely used in data centers with the support of erasure codes. Among many RAID-6 implementations, one set of codes called Maximum Distance Separable (MDS) codes, aim to offer data protection against disk failures with optimal storage efficiency. Since today's large data centers typically shift to the services of cloud computing, a challenging issue is how to accelerate the scaling process of RAID-6 systems based on MDS codes. To address this challenge, we propose a novel MDS Code Scaling Framework (MDS-Frame), which is a unified management scheme on various MDS codes to achieve high scalability. It bridges various MDS codes for flexible scaling via several intermediate codes. In our mathematical analysis, compared to typical RAID scaling approaches, MDS-Frame shows its advantages in the following aspects: It reduces more than 44.1% migration I/Os, saves the migration time by up to 95.2%, and speeds up the migration process by a factor of up to 20.7. Chentao Wu, Xubin He |
ICPP | 1 |
| 2012 | SDM: A Stripe-Based Data Migration Scheme to Improve the Scalability of RAID-6abstractIn large scale data storage systems, RAID-6 has received more attention due to its capability to tolerate concurrent failures of any two disks, providing a higher level of reliability. However, a challenging issue is its scalability, or how to efficiently expand the disks. The main reason causing this problem is the typical fault tolerant scheme of most RAID-6 systems known as Maximum Distance Separable (MDS) codes, which offer data protection against disk failures with optimal storage efficiency but they are difficult to scale. To address this issue, we propose a novel Stripe-based Data Migration (SDM) scheme for large scale storage systems based on RAID-6 to achieve higher scalability. SDM is a stripe-level scheme, and the basic idea of SDM is optimizing data movements according to the future parity layout, which minimizes the overhead of data migration and parity modification. SDM scheme also provides uniform data distribution, fast data addressing and migration. We have conducted extensive mathematical analysis of applying SDM to various popular RAID-6 coding methods such as RDP, P-Code, H-Code, HDP, X-Code, and EVENODD. The results show that, compared to existing scaling approaches, SDM decreases more than 72.7% migration I/O operations and saves the migration time by up to 96.9%, which speeds up the scaling process by a factor of up to 32. Chentao Wu, Xubin He, Jizhong Han, Huailiang Tan, Changsheng Xie 0001 |
CLUSTER | 1 |
| 2012 | GSR: A Global Stripe-Based Redistribution Approach to Accelerate RAID-5 ScalingabstractUnder the severe energy crisis and the fast development of cloud computing, nowadays sustainability in large data centers receives much more attention than ever. Due to its high performance and reliability, RAID, particularly RAID-5, is widely used in these data centers. However, a challenge on the sustainability of RAID-5 is its scalability, or how to efficiently expand/reduce the disks. The main reason causing this problem is the special layout of RAID-5 with parity blocks. To address this problem, in this paper, we propose a novel redistribution approach to accelerate RAID-5 scaling, called Global Stripe-based Redistribution (GSR). The basic idea is to maintain the layout of most stripes while sacrificing a small portion of stripes according to a global view of all stripes. GSR has four main advantages: (1) It supports bi-directional RAID-5 scaling (both scale-up and scale-down), (2) GSR minimizes the overhead of scaling process, including the data migration cost, parity modification and computation cost, and the operations of metadata, (3) Different from previous approaches, GSR provides high flexibility and high availability for the write requests, (4) A disk array can achieve higher capacity, performance and storage efficiency by extending more disks via GSR. In our mathematical analysis, GSR maintains uniform distribution, saves up to 81.5% I/O operations and reduces the data migration time by up to 68.0%, which speeds up the scaling process by a factor of up to 3.13. Chentao Wu, Xubin He |
ICPP | 1 |
| 2011 | HDP code: A Horizontal-Diagonal Parity Code to Optimize I/O load balancing in RAID-6abstractWith higher reliability requirements in clusters and data centers, RAID-6 has gained popularity due to its capability to tolerate concurrent failures of any two disks, which has been shown to be of increasing importance in large scale storage systems. Among various implementations of erasure codes in RAID-6, a typical set of codes known as Maximum Distance Separable (MDS) codes aim to offer data protection against disk failures with optimal storage efficiency. However, because of the limitation of horizontal parity or diagonal/anti-diagonal parities used in MDS codes, storage systems based on RAID-6 suffers from unbalanced I/O and thus low performance and reliability. To address this issue, in this paper, we propose a new parity called Horizontal-Diagonal Parity (HDP), which takes advantages of both horizontal and diagonal/anti-diagonal parities. The corresponding MDS code, called HDP code, distributes parity elements uniformly in each disk to balance the I/O workloads. HDP also achieves high reliability via speeding up the recovery under single or double disk failure. Our analysis shows that HDP provides better balanced I/O and higher reliability compared to other popular MDS codes. Chentao Wu, Xubin He, Guanying Wu, Shenggang Wan, Qiang Cao 0001, Changsheng Xie 0001 |
DSN | 1 |
| 2011 | H-Code: A Hybrid MDS Array Code to Optimize Partial Stripe Writes in RAID-6abstractRAID-6 is widely used to tolerate concurrent failures of any two disks to provide a higher level of reliability with the support of erasure codes. Among many implementations, one class of codes called Maximum Distance Separable (MDS) codes aims to offer data protection against disk failures with optimal storage efficiency. Typical MDS codes contain horizontal and vertical codes. Due to the horizontal parity, in the case of partial stripe write (refers to I/O operations that write new data or update data to a subset of disks in an array) in a row, horizontal codes may get less I/O operations in most cases, but suffer from unbalanced I/O distribution. They also have limitation on high single write complexity. Vertical codes improve single write complexity compared to horizontal codes, while they still suffer from poor performance in partial stripe writes. In this paper, we propose a new XOR-based MDS array code, named Hybrid Code (H-Code), which optimizes partial stripe writes for RAID-6 by taking advantages of both horizontal and vertical codes. H-Code is a solution for an array of (p+1) disks, where p is a prime number. Unlike other codes taking a dedicated anti-diagonal parity strip, H-Code uses a special anti-diagonal parity layout and distributes the anti-diagonal parity elements among disks in the array, which achieves a more balanced I/O distribution. On the other hand, the horizontal parity of H-Code ensures a partial stripe write to continuous data elements in a row share the same row parity chain, which can achieve optimal partial stripe write performance. Not only within a row but also within a stripe, H-Code offers optimal partial stripe write complexity to two continuous data elements and optimal partial stripe write performance among all MDS codes to the best of our knowledge. Specifically, compared to RDP and EVENODD codes, H-Code reduces I/O cost by up to 15.54% and 22.17%. Overall, H-code has optimal storage efficiency, optimal encoding/decoding computational complexity, optimal complexity of both single write and partial stripe write. Chentao Wu, Shenggang Wan, Xubin He, Qiang Cao 0001, Changsheng Xie 0001 |
IPDPS | 1 |
| 2010 | Hint-K: An Efficient Multi-level Cache Using K-Step HintsabstractI/O performance has been critical for large scale distributed systems. Many approaches, including hint-based multi-level cache, have been proposed to smooth the gap between different levels. These solutions demote or promote cache blocks based on the latest history information, which is insufficient for applications where frequent demote and promote operations occur. In this paper we propose a novel multi-level buffer cache using K-step hints (Hint-K) to improve the I/O performance of distributed systems. The basic idea is to promote a block from the lower level cache to the higher level or demote a block vice versa based on the block’s previous K-step promote or demote operations, which are referred to as K-step hints. If we make an analogy between Hint-K and LRU-K, LRU-K keeps track of the times of last K references for blocks within a single cache level, while our Hint-K keeps track of the information of the last K movements (either demote or promote) of blocks among different cache levels. We develop our Hint-K algorithm and design a mathematical model that can efficiently describe the activeness of any blocks in any cache level. Simulation results show that Hint-K achieves better performance compared to current popular multi-level cache schemes such as PROMOTE, DEMOTE, and MQ under different representative I/O workloads. Chentao Wu, Xubin He, Qiang Cao 0001, Changsheng Xie 0001 |
ICPP | 1 |
| 2010 | An Evaluation of Two Typical RAID-6 Codes on Online Single Disk Failure RecoveryabstractRedundant Arrays of Independent Disks RAID is a popular storage architecture with high performance and reliability. RAID-6 with a higher level of reliability based on MDS (Maximum Distance Separable) code is well studied, for its optimal storage efficiency. RAID-6 could offer continuous services in degraded mode, during the period of online failure recovery. However, the online recovery would bring a considerable I/O workflow to the storage system, that almost all the surviving data in the system need to be accessed. Due to the limitation of disk bandwidth, user response time would be significantly affected by the recovery workflow. In this paper, we examine the online recovery performance of two typical MDS RAID-6 codes RDP code and P-code. To our observation, P-code significiantly outperforms RDP in user response time and recovery duration during a single disk failure recovery. To our analysis, the difference comes from not only the parity layout but also the parity organization. Therefore, we propose a new categorization for existing MDS RAID-6 codes, based on the methodology of parity organization. By our approach, all the MDS RAID-6 codes could be categorized to Sym-codes with only one type of parity, and Asym-codes with at least two different types of parity. Qiang Cao 0001, Shenggang Wan, Chentao Wu, Shenghui Zhan |
NAS | 3 |
| 2009 | Hotspot Prediction and cache in distributed stream-processing storage systemsabstractStorage performance is critical in today's distributed stream-processing systems. One approach to improve the performance is to use hotspot attribute in object-based storage systems. This paper discusses hotspot classification and identification, and then presents an object hotspot prediction model (OHPM) to dynamically predict hotspots. Based on this model, we discuss an efficient hotspot caching strategy to improve the performance. To demonstrate the effectiveness of our proposed approach, we have developed a prototype of hotspot attribute-managed storage system (HASS) by extending object-based storage device (OSD) file system and iSCSI protocols. Experimental results show that the HASS improves the throughput by up to 62% and reduces the disk I/O by as much as 25% in our VoD tests by integrating our object hotspot prediction and cache approaches. Chentao Wu, Xubin He, Shenggang Wan, Qiang Cao 0001, Changsheng Xie 0001 |
IPCCC | 1 |
| 2008 | An Adaptive Cache Management Using Dual LRU Stacks to Improve Buffer Cache PerformanceabstractCache plays an essential role in modern computer systems to smooth the performance gap between memory and CPU. Most existing cache replacement algorithms use three stacks: recency stack, frequency stack and history stack. The balance and design of those stacks is a key to achieve high hit ratio, thus improving the buffer cache efficiency. In this paper we propose a new cache replacement algorithm, adaptive dual LRU, or AD-LRU for short, to efficiently utilize the buffer cache pages. Instead of using one LRU stack, we use two LRU stacks: one LRU stack LR to catch the accesses of pages with low recency, and the other LRU stack HR to catch the accesses of pages with high recency. The idea is to adaptively adjust the sizes of the history stack, recency and frequency stacks, an overall buffer cache efficiency in terms of hit ratio will be improved. Simulations results show that AD-LRU demonstrates higher hit ratio compared to existing popular algorithms such as LRU, ARC, and LIRS. Shenggang Wan, Qiang Cao 0001, Xubin He, Changsheng Xie 0001, Chentao Wu |
IPCCC | 5 |