EDBT 2026 Demo / reviewers in the wild / expert
Qingsong Wei
dblp:59/6202
· DBLP profile ↗
47ranked-venue papers
15as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 14 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
15 papers |
Storage systems · 44% Memory systems · 28% Cloud and datacenter computing · 13% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 76% Deep learning architectures and training · 24% | |
| Network and information security
4 papers |
Blockchain and cryptocurrency security · 58% Security and privacy of machine learning · 27% Systems and software security · 15% |
Topics — the 30 heaviest of 54, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
2.4 | 8 | 2020 | NV-Journaling: Locality-Aware Journaling Using Byte-Addressable Non-Volatile Memory · IEEE Trans. Computers 2020 Persisting RB-Tree into NVM in a Consistency Perspective · ACM Trans. Storage 2018 NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018 |
Machine learning › Efficient and distributed learning
federated learning |
1.6 | 2 | 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning · AAAI 2025 An Aggregation-Free Federated Learning for Tackling Data Heterogeneity · CVPR 2024 |
Storage systems
file systems |
1.3 | 5 | 2020 | NV-Journaling: Locality-Aware Journaling Using Byte-Addressable Non-Volatile Memory · IEEE Trans. Computers 2020 Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017 Transactional NVM cache with high performance and crash consistency · SC 2017 |
Storage systems
crash consistency |
1.2 | 5 | 2020 | NV-Journaling: Locality-Aware Journaling Using Byte-Addressable Non-Volatile Memory · IEEE Trans. Computers 2020 Persisting RB-Tree into NVM in a Consistency Perspective · ACM Trans. Storage 2018 Transactional NVM cache with high performance and crash consistency · SC 2017 |
Storage systems › file systems
journaling file system |
1.0 | 3 | 2020 | NV-Journaling: Locality-Aware Journaling Using Byte-Addressable Non-Volatile Memory · IEEE Trans. Computers 2020 Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017 Transactional NVM cache with high performance and crash consistency · SC 2017 |
Blockchain and cryptocurrency security
smart contract security |
1.0 | 1 | 2026 | $AiRacleX$: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation · IEEE Trans. Serv. Comput. 2026 |
Systems and software security
vulnerability discovery |
1.0 | 1 | 2026 | $AiRacleX$: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation · IEEE Trans. Serv. Comput. 2026 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning |
0.9 | 1 | 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning · AAAI 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning · AAAI 2025 |
Machine learning › Efficient and distributed learning › federated learning
personalized federated learning |
0.9 | 1 | 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning · AAAI 2025 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.9 | 1 | 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning · AAAI 2025 |
Machine learning › Deep learning architectures and training
state space model |
0.9 | 1 | 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning · AAAI 2025 |
Blockchain and cryptocurrency security
blockchain interoperability |
0.9 | 1 | 2025 | Atomic Smart Contract Interoperability With High Efficiency via Cross-Chain Integrated Execution · IEEE Trans. Parallel Distributed Syst. 2025 |
Security and privacy of machine learning
federated learning |
0.9 | 1 | 2025 | Maximizing Uncertainty for Federated Learning via Bayesian Optimization-Based Model Poisoning · IEEE Trans. Inf. Forensics Secur. 2025 |
Security and privacy of machine learning › poisoning attack
model poisoning attack |
0.9 | 1 | 2025 | Maximizing Uncertainty for Federated Learning via Bayesian Optimization-Based Model Poisoning · IEEE Trans. Inf. Forensics Secur. 2025 |
Blockchain and cryptocurrency security
smart contract |
0.9 | 1 | 2025 | Atomic Smart Contract Interoperability With High Efficiency via Cross-Chain Integrated Execution · IEEE Trans. Parallel Distributed Syst. 2025 |
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory |
0.8 | 3 | 2020 | NV-Journaling: Locality-Aware Journaling Using Byte-Addressable Non-Volatile Memory · IEEE Trans. Computers 2020 Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017 Transactional NVM cache with high performance and crash consistency · SC 2017 |
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning |
0.8 | 1 | 2024 | An Aggregation-Free Federated Learning for Tackling Data Heterogeneity · CVPR 2024 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.8 | 1 | 2024 | Towards a Heterogeneous and Elastic Cloud Service System With a Correlation-Based Universal Resource Matching Strategy · IEEE Trans. Serv. Comput. 2024 |
Cloud and datacenter computing › resource management › resource allocation and scheduling
resource matching |
0.8 | 1 | 2024 | Towards a Heterogeneous and Elastic Cloud Service System With a Correlation-Based Universal Resource Matching Strategy · IEEE Trans. Serv. Comput. 2024 |
Performance modeling and evaluation › capacity planning
resource requirement estimation |
0.8 | 1 | 2024 | Towards a Heterogeneous and Elastic Cloud Service System With a Correlation-Based Universal Resource Matching Strategy · IEEE Trans. Serv. Comput. 2024 |
Memory systems
cache management |
0.5 | 2 | 2016 | Improving Flash-Based Disk Cache with Lazy Adaptive Replacement · ACM Trans. Storage 2016 Sizing Cleancache Allocation for Virtual Machines' Transcendent Memory · IEEE Trans. Computers 2016 |
Storage systems › data reduction
data deduplication |
0.3 | 1 | 2018 | NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018 |
Storage systems › data reduction › data deduplication
inline deduplication |
0.3 | 1 | 2018 | NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018 |
Storage systems
i/o scheduling |
0.3 | 1 | 2018 | Dynamic Scheduling with Service Curve for QoS Guarantee of Large-Scale Cloud Storage · IEEE Trans. Computers 2018 |
Cloud and datacenter computing
quality of service |
0.3 | 1 | 2018 | Dynamic Scheduling with Service Curve for QoS Guarantee of Large-Scale Cloud Storage · IEEE Trans. Computers 2018 |
Storage systems › file systems
versioning |
0.3 | 1 | 2018 | Persisting RB-Tree into NVM in a Consistency Perspective · ACM Trans. Storage 2018 |
Storage systems › file systems › file system design
persistent memory file system |
0.3 | 2 | 2018 | NV-Tree: Reducing Consistency Cost for NVM-based Single Level Systems · FAST 2015 NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018 |
Distributed systems
consensus |
0.3 | 1 | 2026 | BriDe Arbitrager: Enhancing Arbitrage in Ethereum 2.0 via Bribery-Enabled Delayed Block Production · IEEE Trans. Dependable Secur. Comput. 2026 |
Distributed systems › consensus › permissionless consensus
proof-of-stake |
0.3 | 1 | 2026 | BriDe Arbitrager: Enhancing Arbitrage in Ethereum 2.0 via Bribery-Enabled Delayed Block Production · IEEE Trans. Dependable Secur. Comput. 2026 |
Methods — techniques the papers use, named apart from their topics
transaction ordering · 2.0game-theoretic bribery strategy · 2.0transaction aggregation · 1.7fine-grained state lock · 1.72PC-based mechanism · 1.7large language model · 1.0knowledge mining · 1.0chain-of-thought prompting · 1.0selective state space model · 0.9bayesian optimization · 0.9adapter tuning · 0.9KL divergence · 0.9soft labels · 0.8peer knowledge sharing · 0.8decision tree · 0.8dataset condensation · 0.8correlation analysis · 0.8locality-aware checkpointing · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-vehicle cooperative decision-making for ramp merging in mixed traffic: A pre-trial behavior-informed reinforcement learning approach
Qingsong Wei, Xianyi Xie, Lisheng Jin, Yewei Shi, Yaping Liao |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | CoFeatNet: An Efficient Multimodal Feature Extraction Network for Cooperative Vehicle-to-Infrastructure 3-D Object Detection
Baicang Guo, Yewei Shi, Lisheng Jin, Qingsong Wei, Jiaguo Liu |
IEEE Internet Things J. | 6 |
| 2026 | BriDe Arbitrager: Enhancing Arbitrage in Ethereum 2.0 via Bribery-Enabled Delayed Block ProductionabstractThe advent of Ethereum 2.0 has introduced significant changes, particularly the shift to Proof-of-Stake consensus. This change presents new opportunities and challenges for arbitrage. Amidst these changes, we introduce BriDe Arbitrager, a novel tool designed for Ethereum 2.0 that leveragesBribery-driven attacks toDelay block production and increase arbitrage gains. The main idea is to allow malicious proposers to delay block production by bribing validators/proposers, thereby gaining more time to identify arbitrage opportunities. Through analysing the bribery process, we design an adaptive bribery strategy. Additionally, we propose a Delayed Transaction Ordering Algorithm to leverage the delayed time to amplify arbitrage profits for malicious proposers. To ensure fairness and automate the bribery process, we design and implement a bribery smart contract and a bribery client. As a result, BriDe Arbitrager enables adversaries controlling a limited ($\lt 1/4$) fraction of the voting powers to delay block production via bribery and arbitrage more profit. Extensive experimental results based on Ethereum historical transactions demonstrate that BriDe Arbitrager yields an average of 8.78 ETH (16,687.88 USD) daily profits. Furthermore, our approach does not trigger any slashing mechanisms and remains effective even under Proposer Builder Separation and other potential mechanisms will be adopted by Ethereum. Hulin Yang, Jin Zhang 0001, Alia Asheralieva, Qingsong Wei, Rick Siow Mong Goh |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Improving Learning of New Diseases Through Knowledge-Enhanced Initialization for Federated Adapter TuningabstractIn healthcare, federated learning (FL) is a widely adopted framework that enables privacy-preserving collaboration among medical institutions. With large foundation models (FMs) demonstrating impressive capabilities, using FMs in FL through cost-efficient adapter tuning has become a popular approach. Given the rapidly evolving healthcare environment, it is crucial for individual clients to quickly adapt to new tasks or diseases by tuning adapters while drawing upon past experiences. In this work, we introduce Federated Knowledge-Enhanced Initialization (FedKEI), a novel framework that leverages cross-client and cross-task transfer from past knowledge to generate informed initializations for learning new tasks with adapters. FedKEI begins with a global clustering process at the server to generalize knowledge across tasks, followed by the optimization of aggregation weights across clusters (inter-cluster weights) and within each cluster (intra-cluster weights) to personalize knowledge transfer for each new task. To facilitate more effective learning of the inter- and intra-cluster weights, we adopt a bi-level optimization scheme that collaboratively learns the global intra-cluster weights across clients and optimizes the local inter-cluster weights toward each client's task objective. Extensive experiments on three benchmark datasets of different modalities, including dermatology, chest X-rays, and retinal OCT, demonstrate FedKEI's advantage in adapting to new diseases compared to state-of-the-art methods. Danni Peng, Yuan Wang 0008, Kangning Cai, Peiyan Ning, Jiming Xu, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei, Huazhu Fu |
IEEE Trans. Medical Imaging | 8 |
| 2026 | $AiRacleX$: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt GenerationabstractDecentralized finance (DeFi) applications depend on accurate price oracles to ensure secure and fair transactions. However, poorly integrated oracles remain susceptible to manipulation, enabling attackers to exploit smart contract logic for unfair asset valuation and financial gain. While many such vulnerabilities are only detected after deployment, smart contracts are typically immutable once deployed, making post-hoc fixes costly or infeasible. This highlights the critical need for detecting oracle manipulation risks before deployment. In this paper, we propose$AiRacleX$, a novel LLM-driven framework that enables pre-deployment detection of price oracle manipulation vulnerabilities by leveraging the complementary strengths of multiple large language models (LLMs). Our approach begins with domain-specific knowledge extraction, where an LLM model synthesizes precise insights about price oracle vulnerabilities, eliminating the need for profound expertise from developers or auditors. This knowledge forms the foundation for a second LLM model to generate structured, context-aware Chain-of-Thought prompts, which guide a third LLM model in accurately identifying manipulation patterns in smart contracts. We evaluate$AiRacleX$on 60 known vulnerabilities from 44 real-world DeFi exploits and Code4rena projects spanning 2021-2023. The results show that$AiRacleX$achieves a 2.58 times improvement in recall over the state-of-the-art GPTScan, with comparable precision. Our framework also demonstrates strong extensibility and efficiency, and supports deployment with open-source LLMs to enhance security and reduce operational cost. Yuan Wang 0008, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh, David Lo 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter TuningabstractPersonalized federated learning (PFL) studies effective model personalization to address the data heterogeneity issue among clients in traditional federated learning (FL). Existing PFL approaches mainly generate personalized models by relying solely on the clients' latest updated models while ignoring their previous updates, which may result in suboptimal personalized model learning. To bridge this gap, we propose a novel framework termed pFedSeq, designed for personalizing adapters to fine-tune a foundation model in FL. In pFedSeq, the server maintains and trains a sequential learner, which processes a sequence of past adapter updates from clients and generates calibrations for personalized adapters. To effectively capture the cross-client and cross-step relations hidden in previous updates and generate high-performing personalized adapters, pFedSeq adopts the powerful selective state space model (SSM) as the architecture of sequential learner. Through extensive experiments on four public benchmark datasets, we demonstrate the superiority of pFedSeq over state-of-the-art PFL methods. Danni Peng, Yuan Wang 0008, Huazhu Fu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei |
AAAI | 7 |
| 2025 | History-Aware and Dynamic Client Contribution in Federated LearningabstractFederated Learning (FL) is a collaborative machine learning (ML) approach, where multiple clients participate in training an ML model without exposing their private data. Fair and accurate assessment of client contributions facilitates incentive allocation in FL and encourages diverse clients to participate in a unified model training. Existing methods for contribution assessment adopts a co-operative game-theoretic concept, called Shapley value, but under restricted assumptions, e.g., all clients’ participating in all epochs or at least in one epoch of FL. We propose a history-aware client contribution assessment framework, called FLContrib, where client-participation is dynamic, i.e., a subset of clients participates in each epoch. The theoretical underpinning of FLContrib is based on the Markovian training process of FL. Under this setting, we directly apply the linearity property of Shapley value and compute a historical timeline of client contributions. Considering the possibility of a limited computational budget, we propose a two-sided fairness criteria to schedule Shapley value computation in a subset of epochs. Empirically, FLContrib is efficient and consistently accurate in estimating contribution across multiple utility functions. As a practical application, we apply FLContrib to detect dishonest clients in FL based on historical Shaplee values. Bishwamittra Ghosh, Debabrota Basu, Huazhu Fu, Yuan Wang 0008, Renuga Kanagavelu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei |
ECAI | 9 |
| 2025 | Maximizing Uncertainty for Federated Learning via Bayesian Optimization-Based Model PoisoningabstractAs we transition from Narrow Artificial Intelligence towards Artificial Super Intelligence, users are increasingly concerned about their privacy and the trustworthiness of machine learning (ML) technology. A common denominator for the metrics of trustworthiness is the quantification of uncertainty inherent in DL algorithms, and specifically in the model parameters, input data, and model predictions. One of the common approaches to address privacy-related issues in DL is to adopt distributed learning such as federated learning (FL), where private raw data is not shared among users. Despite the privacy-preserving mechanisms in FL, it still faces challenges in trustworthiness. Specifically, the malicious users, during training, can systematically create malicious model parameters to compromise the models’ predictive and generative capabilities, resulting in high uncertainty about their reliability. To demonstrate malicious behaviour, we propose a novel model poisoning attack method named Delphi which aims to maximise the uncertainty of the global model output. We achieve this by taking advantage of the relationship between the uncertainty and the model parameters of the first hidden layer of the local model. Delphi employs two types of optimisation, Bayesian Optimisation and Least Squares Trust Region, to search for the optimal poisoned model parameters, named as Delphi-BO and Delphi-LSTR. We quantify the uncertainty using the KL Divergence to minimise the distance of the predictive probability distribution towards an uncertain distribution of model output. Furthermore, we establish a mathematical proof for the attack effectiveness demonstrated in FL. Numerical results demonstrate that Delphi-BO induces a higher amount of uncertainty than Delphi-LSTR highlighting vulnerability of FL systems to model poisoning attacks. Marios Aristodemou, Xiaolan Liu 0001, Yuan Wang 0008, Konstantinos G. Kyriakopoulos, Sangarapillai Lambotharan, Qingsong Wei |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Atomic Smart Contract Interoperability With High Efficiency via Cross-Chain Integrated ExecutionabstractWith the development of Ethereum, numerous blockchains compatible with Ethereum's execution environment (i.e., Ethereum Virtual Machine, EVM) have emerged. Developers can leverage smart contracts to run various complex decentralized applications on top of blockchains. However, the increasing number of EVM-compatible blockchains has introduced significant challenges in cross-chain interoperability, particularly in ensuring efficiency and atomicity for the whole cross-chain application. Existing solutions areeither limited in guaranteeing overall atomicity for the cross-chain application, or inefficient due to the need for multiple rounds of cross-chain smart contract execution.To address this gap, we proposeIntegrateX, an efficient cross-chain interoperability system that ensures the overall atomicity of cross-chain smart contract invocations. The core idea is todeploy the logic required for cross-chain execution onto a single blockchain, where it can be executed in an integrated manner.This allows cross-chain applications to perform all cross-chain logic efficiently within the same blockchain.IntegrateXconsists of across-chain smart contract deployment protocoland across-chain smart contract integrated execution protocol.The former achieves efficient and secure cross-chain deployment by decoupling smart contract logic from state, and employing an off-chain cross-chain deployment mechanism combined with on-chain cross-chain verification. The latter ensures atomicity of cross-chain invocations through a 2PC-based mechanism, and enhances performance through transaction aggregation and fine-grained state lock. We implement a prototype ofIntegrateX. Extensive experiments demonstrate that it reduces up to 61.2% latency compared to the state-of-the-art baseline while maintaining low gas consumption. Chaoyue Yin, Jin Zhang 0001, You Lin, Qingsong Wei, Rick Siow Mong Goh |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | An Aggregation-Free Federated Learning for Tackling Data HeterogeneityabstractThe performance of Federated Learning (FL) hinges on the effectiveness of utilizing knowledge from distributed datasets. Traditional FL methods adopt an aggregate-then-adapt framework, where clients update local models based on a global model aggregated by the server from the previous training round. This process can cause client drift, especially with significant cross-client data heterogeneity, impacting model performance and convergence of the FL algorithm. To address these challenges, we introduce FedAF, a novel aggregation-free FL algorithm. In this framework, clients collaboratively learn condensed data by leveraging peer knowledge, the server subsequently trains the global model using the condensed data and soft labels received from the clients. FedAF inherently avoids the issue of client drift, enhances the quality of condensed data amid notable data heterogeneity, and improves the global model performance. Extensive numerical studies on several popular benchmark datasets show FedAF surpasses various state-of-the-art FL algorithms in handling label-skew and feature-skew data heterogeneity, leading to superior global model accuracy and faster convergence. Yuan Wang 0008, Huazhu Fu, Renuga Kanagavelu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh |
CVPR | 4 |
| 2024 | MedSynth: Leveraging Generative Model for Healthcare Data Sharing
Renuga Kanagavelu, Madhav Walia, Yuan Wang 0008, Huazhu Fu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh |
MICCAI (12) | 5 |
| 2024 | Towards a Heterogeneous and Elastic Cloud Service System With a Correlation-Based Universal Resource Matching StrategyabstractIn elastic cloud service systems, it is a challenge to evaluate and match the fluctuating resource demand of workloads. Existing studies typically monitor workload characteristics and build models that map these characteristics to actual demand. However, workload characteristics are multidimensional, and the impact of each dimension on resource demand differs, so it requires differentiated treatment when building models. This paper proposes a Correlation-Based Universal Resource Matching (CBURM) strategy to realize a Heterogeneous and Elastic Cloud Service System (HECSS). CBURM consists of a Correlation-based resource Demand Evaluation (CDE) method and a Universal Resource Measurement (URM) scheme. Specifically, CDE discriminates the relevance of each dimension in workload characteristics, based on the correlations between workload characteristics and the demand. Then, it generates resource demand decisions dimension by dimension, from the most relevant to the least relevant dimensions. After that, it generates a complete decision tree model to evaluate subsequent workload demand for heterogeneous resources. Finally, URM optimizes the resource allocation to achieve a low-overhead resource matching. Experimental results show that, URM reduces the total comprehensive operation cost by 82%+, compared to a normal resource allocation scheme. Additionally, CDE outperforms two state-of-the-art methods (LTP and 2SP), with its performance closer to the ideal baseline. Specifically, CDE achieves a 40.275% overall resource saving rate, which is 38.62% higher than LTP and 8.46% higher than 2SP. Besides, CDE achieves a 92.43% average service quality satisfaction ratio, higher than the 82.9% and 88.83% achieved respectively by LTP and 2SP. Cheng Hu 0004, Yuhui Deng 0001, Wenyu Luo, Qingsong Wei, Geyong Min |
IEEE Trans. Serv. Comput. | 4 |
| 2020 | NV-Journaling: Locality-Aware Journaling Using Byte-Addressable Non-Volatile MemoryabstractModern file systems rely on the journaling mechanism to maintain crash consistency. The use of non-volatile memory (NVM) significantly improves the performance of journaling file systems. However, the superior performance of NVM will increase the likelihood of the journal filling up more often, thereby increasing the frequency of checkpointing. Together with the large amount of random checkpointing I/O found in most use cases, the checkpointing process becomes a new performance bottleneck. This paper proposes NV-Journaling, a strategy that reduces the frequency of checkpointing as well as reshapes the I/O pattern of checkpointing from one of random I/O to that which is more sequential I/O. NV-Journaling introduces fine-grained commits along with a cache-friendly NVM journaling layout that exploits the idiosyncrasies of NVM technology. Under this scheme, only the modified portion of a block, rather than the entire block, is written into the NVM journal device. Doing so significantly reduces checkpoint frequency and achieves better space utilization. NV-Journaling further reshapes the I/O pattern of checkpoint using a locality-aware checkpointing process. Checkpointed blocks are classified into hot and cold blocks. NV-Journaling maintains a hot block list to absorb repeated updates, and a cold bucket list to group blocks by their proximity on disk. When a checkpoint is required, cold buckets are selected such that blocks are sequentially flushed to the hard disk. We built a prototype of NV-Journaling by modifying the JBD2 layer in the Linux kernel and evaluated it using different workloads. Our experimental results show that NV-Journaling can improve performance by up to 4.3× compared to traditional journaling. Cheng Chen 0008, Qingsong Wei, Weng-Fai Wong, Chundong Wang 0001 |
IEEE Trans. Computers | 2 |
| 2018 | NV-Dedup: High-Performance Inline Deduplication for Non-Volatile MemoryabstractThe byte-addressable non-volatile memory (NVM) is a promising medium for data storage. NVM-oriented file systems have been designed to explore NVM's performance potential. Meanwhile, applications may write considerable duplicate data. For NVM, a removal of duplicate data can promote space efficiency, improve write endurance, and potentially improve the performance by avoidance of repeatedly writing the same data. However, we have observed severe performance degradations when implementing a state-of-the-art inline deduplication algorithm in an NVM-oriented file system. A quantitative analysis reveals that, with NVM, 1) the conventional way to manage deduplication metadata for block devices, particularly in light of consistency, is inefficient, and, 2) the performance with deduplication becomes more subject to fingerprint calculations. We hence propose a deduplication algorithm called NV-Dedup. NV-Dedup manages deduplication metadata in a fine-grained, CPU and NVM-favored way, and preserves the metadata consistency with a lightweight transactional scheme. It also does workload-adaptive fingerprinting based on an analytical model and a transition scheme among fingerprinting methods to reduce calculation penalties. We have built a prototype of NV-Dedup in the Persistent Memory File System (PMFS). Experiments show that, NV-Dedup not only substantially saves NVM space, but also boosts the performance of PMFS by up to 2.1x. Chundong Wang 0001, Qingsong Wei, Jun Yang 0022, Cheng Chen 0008, Yechao Yang, Mingdi Xue |
IEEE Trans. Computers | 2 |
| 2018 | Dynamic Scheduling with Service Curve for QoS Guarantee of Large-Scale Cloud StorageabstractWith the growing popularity of cloud storage, more and more diverse applications with diverse service level agreements (SLAs) are being accommodated into it. The quality of service (QoS) support for applications in a shared cloud storage becomes important. However, performance isolation, diverse performance requirements, especially harsh latency guarantees and high system utilization, are all challenging and desirable for QoS design. In this paper, we propose a service curve-based QoS algorithm to support latency guarantee applications, IOPS guarantee applications and best-effort applications at the same storage system, which not only provides a QoS guarantee for applications, but also pursues better system utilization. Three priority queues are exploited and different service curves are applied for different types of applications. I/O requests from different applications are scheduled and dispatched among the three queues according to their service curves and I/O urgency status, so that QoS requirements of all applications can be guaranteed on the shared storage system. Our experimental results show that our algorithm not only simultaneously guarantees the QoS targets of latency and throughput (IOPS), but also improves the utilization of storage resources. Yu Zhang 0028, Qingsong Wei, Cheng Chen 0008, Mingdi Xue, Xinkun Yuan, Chundong Wang 0001 |
IEEE Trans. Computers | 2 |
| 2018 | Persisting RB-Tree into NVM in a Consistency PerspectiveabstractByte-addressable non-volatile memory (NVM) is going to reshape conventional computer systems. With advantages of low latency, byte-addressability, and non-volatility, NVM can be directly put on the memory bus to replace DRAM. As a result, both system and application softwares have to be adjusted to perceive the fact that the persistent layer moves up to the memory. However, most of the current in-memory data structures will be problematic with consistency issues if not well tuned with NVM. This article places emphasis on an important in-memory structure that is widely used in computer systems, i.e., the Red/Black-tree (RB-tree). Since it has a long and complicated update process, the RB-tree is prone to inconsistency problems with NVM. This article presents an NVM-compatible consistent RB-tree with a new technique named cascade-versioning . The proposed RB-tree (i) is all-time consistent and scalable and (ii) needs no recovery procedure after system crashes. Experiment results show that the RB-tree for NVM not only achieves the aim of consistency with insignificant spatial overhead but also yields comparable performance to an ordinary volatile RB-tree. Chundong Wang 0001, Qingsong Wei, Lingkun Wu, Sibo Wang 0001, Cheng Chen 0008, Xiaokui Xiao, Jun Yang 0022, Mingdi Xue, Yechao Yang |
ACM Trans. Storage | 2 |
| 2017 | Transactional NVM cache with high performance and crash consistencyabstractThe byte-addressable non-volatile memory (NVM) is new promising storage medium. Compared to NAND flash memory, the next-generation NVM not only preserves the durability of stored data but has much shorter access latencies. An architect can utilize the fast and persistent NVM as an external disk cache. Regarding the system's crash consistency, a prevalent journaling file system needs to run atop an NVM disk cache. However, the performance is severely impaired by redundant efforts in achieving crash consistency in both file system and disk cache. Therefore, we propose a new mechanism called transactional NVM disk cache (Tinca). In brief, Tinca jointly guarantees consistency of file system and disk cache and removes the performance penalty of file system journaling with a lightweight transaction scheme. Evaluations confirm that Tinca significantly outperforms state-of-the-art design by up to 2.5X in local and cluster tests without causing any inconsistency issue. Qingsong Wei, Chundong Wang 0001, Cheng Chen 0008, Yechao Yang, Jun Yang 0022, Mingdi Xue |
SC | 1 |
| 2017 | Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVMabstractJournaling file systems have been widely adopted to support applications that demand data consistency. However, we observed that the overhead of journaling can cause up to 48.2% performance drop under certain kinds of workloads. On the other hand, the emerging high-performance, byte-addressable Non-volatile Memory (NVM) has the potential to minimize such overhead by being used as the journal device. The traditional journaling mechanism based on block devices is nevertheless unsuitable for NVM due to the write amplification of metadata journal we observed. In this article, we propose a fine-grained metadata journal mechanism to fully utilize the low-latency byte-addressable NVM so that the overhead of journaling can be significantly reduced. Based on the observation that conventional block-based metadata journal contains up to 90% clean metadata that is unnecessary to be journalled, we design a fine-grained journal format for byte-addressable NVM which contains only modified metadata. Moreover, we redesign the process of transaction committing, checkpointing, and recovery in journaling file systems utilizing the new journal format. Therefore, thanks to the reduced amount of ordered writes for journals, the overhead of journaling can be reduced without compromising the file system consistency. To evaluate our fine-grained metadata journaling mechanism, we have implemented a journaling file system prototype based on Ext4 and JBD2 in Linux. Experimental results show that our NVM-based fine-grained metadata journaling is up to 15.8 × faster than the traditional approach under FileBench workloads. Cheng Chen 0008, Jun Yang 0022, Qingsong Wei, Chundong Wang 0001, Mingdi Xue |
ACM Trans. Storage | 3 |
| 2016 | Extending SSD Lifetime with Persistent In-Memory Metadata ManagementabstractFlash-based solid state drive (SSD) is now widely deployed to speed up data intensive applications. However, I/O amplifications caused by file system metadata and journaling shorten the lifetime of SSD. In this paper, a mechanism named Persistent In-memory Metadata Management (referred to as PIMM) is proposed to reduce I/O traffics to SSD by exploiting the persistency and byte-addressability of Non-volatile Memory (NVM). The PIMM decouples data and metadata access paths, putting data on SSD and metadata in NVM at runtime. Thus, metadata is accessed in byte-addressable manner via the memory bus and metadata I/O is eliminated because metadata in NVM is not flushed back to SSD anymore. The PIMM is prototyped on real NVDIMM platform. Extensive evaluations on implemented prototype show that the proposed PIMM reduces the block erase for SSD by up to 91% and improves performance for different workloads. Qingsong Wei, Cheng Chen 0008, Mingdi Xue, Chundong Wang 0001, Jun Yang 0022 |
CLUSTER | 1 |
| 2016 | Fine-grained metadata journaling on NVMabstractJournaling file systems have been widely used where data consistency must be assured. However, we observed that the overhead of journaling can cause up to 48.2% performance drop under certain kinds of workloads. On the other hand, the emerging high-performance, byte-addressable Non-volatile Memory (NVM) has the potential to minimize such overhead by being used as the journal device. The traditional journaling mechanism based on block devices is nevertheless unsuitable for NVM due to the write amplification of metadata journal we observed. In this paper, we propose a fine-grained metadata journal mechanism to fully utilize the low-latency byte-addressable NVM so that the overhead of journaling can be significantly reduced. Based on the observation that conventional block-based metadata journal contains up to 90% clean metadata that is unnecessary to be journalled, we design a fine-grained journal format for byte-addressable NVM which contains only modified metadata. Moreover, we redesign the process of transaction committing, checkpointing and recovery in journaling file systems utilizing the new journal format. Therefore, thanks to the reduced amount of ordered writes to NVM, the overhead of journaling can be reduced without compromising the file system consistency. Experimental results show that our NVM-based fine-grained metadata journaling is up to 15.8× faster than the traditional approach under FileBench workloads. Cheng Chen 0008, Jun Yang 0022, Qingsong Wei, Chundong Wang 0001, Mingdi Xue |
MSST | 3 |
| 2016 | Sizing Cleancache Allocation for Virtual Machines' Transcendent MemoryabstractThe old virtualization idea and the new multicore technology now make it possible to consolidate multiple workloads on one physical host. This helps reduce the amount of idle resources. In particular, transcendent memory is a recent idea to gather idle memory into a pool that is shared by virtual machines (VMs). It can be viewed as a new level in the memory hierarchy, between main memory and disks. Cleancache is the part of transcendent memory that is used for caching the VMs' clean pages. This paper shows that a Cache Miss Equation can accurately capture observed cleancache behavior, despite its unusual design. The equation can be used to dynamically size and partition cleancache for the VMs. Experiments with a variety of workloads show that this equation-based allocation can help to drastically reduce disk reads and fairly allocate cleancache space. Vimalraj Venkatesan, Y. C. Tay, Qingsong Wei |
IEEE Trans. Computers | 3 |
| 2016 | NV-Tree: A Consistent and Workload-Adaptive Tree Structure for Non-Volatile MemoryabstractThe non-volatile memory (NVM) which can provide DRAM-like performance and disk-like persistency has the potential to build single-level systems by replacing both DRAM and disk. Keeping data consistency in such systems is non-trivial because memory writes may be reordered by CPU. Although ordered memory writes for achieving data consistency can be implemented using the memory fence and the CPU cache line flush instructions, they introduce a significant overhead (more than 10X slower in performance). In this paper, we focus on an important and common data structure, B$^+$Tree. Based on our quantitative analysis for consistent tree structures, we propose NV-Tree, a consistent, cache-optimized and workload-adaptive B$^+$Tree variant with significantly reduced consistency cost (up to 96 percent reduction in CPU cache line flush). To further optimize NV-Tree under various workloads, we propose a workload-adaptive scheme in which the sizes of individual nodes can be dynamically adjusted to improve the performance over time. We implement and evaluate NV-Tree and NV-Store, a key-value store based on NV-Tree, on an NVDIMM server. NV-Tree outperforms the state-of-art consistent tree structures by up to 12X under write-intensive workloads. NV-Store increases the throughput by up to 7.3X under YCSB workloads compared to Redis. Jun Yang 0022, Qingsong Wei, Chundong Wang 0001, Cheng Chen 0008, Khai Leong Yong, Bingsheng He |
IEEE Trans. Computers | 2 |
| 2016 | Improving Flash-Based Disk Cache with Lazy Adaptive ReplacementabstractFor years, the increasing popularity of flash memory has been changing storage systems. Flash-based solid-state drives (SSDs) are widely used as a new cache tier on top of hard disk drives (HDDs) to speed up data-intensive applications. However, the endurance problem of flash memory remains a concern and is getting worse with the adoption of MLC and TLC flash. In this article, we propose a novel cache management algorithm for flash-based disk cache named Lazy Adaptive Replacement Cache (LARC). LARC adopts the idea of selective caching to filter out seldom accessed blocks and prevent them from entering cache. This avoids cache pollution and preserves popular blocks in cache for a longer period of time, leading to a higher hit rate. Meanwhile, by avoiding unnecessary cache replacements, LARC reduces the volume of data written to the SSD and yields an SSD-friendly access pattern. In this way, LARC improves the performance and endurance of the SSD at the same time. LARC is self-tuning and incurs little overhead. It has been extensively evaluated by both trace-driven simulations and synthetic benchmarks on a prototype implementation. Our experiments show that LARC outperforms state-of-art algorithms for different kinds of workloads and extends SSD lifetime by up to 15.7 times. Sai Huang, Qingsong Wei, Dan Feng 0001, Jianxi Chen, Cheng Chen 0008 |
ACM Trans. Storage | 2 |
| 2015 | Accelerating non-volatile/hybrid processor cache design space exploration for application specific embedded systemsabstractIn this article, we propose a technique to accelerate nonvolatile or hybrid of volatile and nonvolatile processor cache design space exploration for application specific embedded systems. Utilizing a novel cache behavior modeling equation and a new accurate cache miss prediction mechanism, our proposed technique can accelerate NVM or hybrid FIFO processor cache design space exploration for SPEC CPU 2000 applications up to 249 times compared to the conventional approach. Mohammad Shihabul Haque, Ang Li 0006, Akash Kumar 0001, Qingsong Wei |
ASP-DAC | 4 |
| 2015 | NV-Tree: Reducing Consistency Cost for NVM-based Single Level Systems
Jun Yang 0022, Qingsong Wei, Cheng Chen 0008, Chundong Wang 0001, Khai Leong Yong, Bingsheng He |
FAST | 2 |
| 2015 | Accelerating Cloud Storage System with Byte-Addressable Non-Volatile MemoryabstractAs building block for cloud storage, distributed file system uses underlying local file systems to manage objects. However, the underlying file system, which is limited by metadata and journaling I/O, significantly affects the performance of the distributed file system. This paper presents an NVM-based file system (referred to as NV-Booster) to accelerate object access for storage node. The NV-Booster leverages byte-addressability and persistency of nonvolatile memory (NVM) to speedup metadata accesses and file system journaling. With NV-Booster, metadata is kept in NVM and accessed in byte-addressable manner through memory bus, while object is stored on hard disk and accessed from I/O bus. In addition, proposed NV-Booster enables fast object search and mapping between object ID and on-disk location with an efficient in-memory namespace management. NV-Booster is implemented in kernel space with NVDIMM and has been extensively evaluated under various workloads. Our experiments show that NV-Booster improves Ceph performance up to 10X, compared to the Ceph with existing local file systems. Qingsong Wei, Mingdi Xue, Jun Yang 0022, Chundong Wang 0001, Cheng Chen 0008 |
ICPADS | 1 |
| 2015 | How to be consistent with persistent memory? An evaluation approachabstractThe advent of the byte-addressable, non-volatile memory (NVM) has initiated the design of new data management strategies to utilize it as the persistent memory (PM). One way to manage the PM is via an in-memory file system. The consistency of the in-memory file system may nevertheless be compromised from directly exposing the PM to the CPU, because data are likely to be flushed from the CPU cache to the PM in an order that is different from the order in which they have been programed to be. As a result, in spite of classic consistency mechanisms, such as journaling and Copy-on-Write, file systems for the PM have to seek support of cacheline flush and memory fence instructions, e.g., clflush and sfence, to achieve ordered writes. On the other hand, manipulating the PM as a consistent block device with conventional file systems is also doable. The pros and cons of two approaches, however, have not been thoroughly investigated yet. We hence do so with extensive evaluations and detailed analyses. Our aim of this paper is to inspire how the PM shall be managed, especially from the performance perspective. Chundong Wang 0001, Qingsong Wei, Jun Yang 0022, Cheng Chen 0008, Mingdi Xue |
NAS | 2 |
| 2015 | Accelerating File System Metadata Access with Byte-Addressable Nonvolatile MemoryabstractFile system performance is dominated by small and frequent metadata access. Metadata is stored as blocks on the hard disk drive. Partial metadata update results in whole-block read or write, which significantly amplifies disk I/O. Furthermore, a huge performance gap between the CPU and disk aggravates this problem. In this article, a file system metadata accelerator (referred to as FSMAC) is proposed to optimize metadata access by efficiently exploiting the persistency and byte-addressability of Nonvolatile Memory (NVM). The FSMAC decouples data and metadata access path, putting data on disk and metadata in byte-addressable NVM at runtime. Thus, data is accessed in a block from I/O the bus and metadata is accessed in a byte-addressable manner from the memory bus. Metadata access is significantly accelerated and metadata I/O is eliminated because metadata in NVM is no longer flushed back to the disk periodically. A lightweight consistency mechanism combining fine-grained versioning and transaction is introduced in the FSMAC. The FSMAC is implemented on a real NVDIMM platform and intensively evaluated under different workloads. Evaluation results show that the FSMAC accelerates the file system up to 49.2 times for synchronized I/O and 7.22 times for asynchronized I/O. Moreover, it can achieve significant performance speedup in network storage and database environment, especially for metadata-intensive or write-dominated workloads. Qingsong Wei, Jianxi Chen, Cheng Chen 0008 |
ACM Trans. Storage | 1 |
| 2015 | Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State DriveabstractExisting space management and address mapping schemes for flash-based Solid-State-Drive (SSD) operate either at page or block granularity, with inevitable limitations in terms of memory requirement, performance, garbage collection, and scalability. To overcome these limitations, we proposed a novel space management and address mapping scheme for flash referred to as Z-MAP, which manages flash space at granularity of Zone. Each Zone consists of multiple numbers of flash blocks. Leveraging workload classification, Z-MAP explores Page-mapping Zone (Page Zone) to store random data and handle a large number of partial updates, and Block-mapping Zone (Block Zone) to store sequential data and lower the overall mapping table. Zones are dynamically allocated and a mapping scheme for a Zone is determined only when it is allocated. Z-MAP uses a small part of Flash memory or phase change memory as a streaming Buffer Zone to log data sequentially and migrate data into Page Zone or Block Zone based on workload classification. A two-level address mapping is designed to reduce the overall mapping table and address translation latency. Z-MAP classifies data before it is permanently stored into Flash memory so that different workloads can be isolated and garbage collection overhead can be minimized. Z-MAP has been extensively evaluated by trace-driven simulation and a prototype implementation on OpenSSD. Our benchmark results conclusively demonstrate that Z-MAP can achieve up to 76% performance improvement, 81% mapping table reduction, and 88% garbage collection overhead reduction compared to existing Flash Translation Layer (FTL) schemes. Qingsong Wei, Cheng Chen 0008, Mingdi Xue, Jun Yang 0022 |
ACM Trans. Storage | 1 |
| 2014 | A 3-Level Cache Miss Model for a Nonvolatile Extension to Transcendent MemoryabstractResource allocation is fundamental to cloud computing, where the memory hierarchy is deep. Space allocation in this hierarchy calls for a model to determine how provisioning at one level affects performance at a lower level. This paper presents a 3-level model that relates the Miss Ratio Curves for two caches at adjacent levels. The model is tested with NEXTmem, which is a transcendent memory used by a Xen hypervisor to cache pages for virtual machines. NEXTmem has a DRAM level and a nonvolatile memory level. The test runs DaCapo benchmarks and shows that the model can be used to enforce fairness at one level, and latency bounds at another level. Vimalraj Venkatesan, Y. C. Tay, Yi Irvette Zhang, Qingsong Wei |
CloudCom | 4 |
| 2014 | CBM: A cooperative buffer management for SSDabstractRandom writes significantly limit the application of Solid State Drive (SSD) in the I/O intensive applications such as scientific computing, Web services, and database. While several buffer management algorithms are proposed to reduce random writes, their ability to deal with workloads mixed with sequential and random accesses is limited. In this paper, we propose a cooperative buffer management scheme referred to as CBM, which coordinates write buffer and read cache to fully exploit temporal and spatial localities among I/O intensive workload. To improve both buffer hit rate and destage sequentiality, CBM divides write buffer space into Page Region and Block Region. Randomly written data is put in the Page Region at page granularity, while sequentially written data is stored in the Block Region at block granularity. CBM leverages threshold-based migration to dynamically classify random write from sequential writes. When a block is evicted from write buffer, CBM merges the dirty pages in write buffer and the clean pages in read cache belonging to the evicted block to maximize the possibility of forming full block write. CBM has been extensively evaluated with simulation and real implementation on OpenSSD. Our testing results conclusively demonstrate that CBM can achieve up to 84% performance improvement and 85% garbage collection overhead reduction compared to existing buffer management schemes. Qingsong Wei, Cheng Chen 0008, Jun Yang 0022 |
MSST | 1 |
| 2014 | Space4time: Optimization latency-sensitive content service in cloud
Lingfang Zeng, Bharadwaj Veeravalli, Qingsong Wei |
J. Netw. Comput. Appl. | 3 |
| 2013 | FSMAC: A file system metadata accelerator with non-volatile memoryabstractFile system performance is dominated by metadata access because it is small and popular. Metadata is stored as block in the file system. Partial metadata update results in whole block read and write which amplifies disk I/O. Huge performance gap between CPU and disk aggravates this problem. In this paper, a file system metadata accelerator (referred as FSMAC) is proposed to optimize metadata access by efficiently exploiting the advantages of Nonvolatile Memory (NVM). FSMAC decouples data and metadata I/O path, putting data on disk and metadata on NVM at runtime. Thus, data is accessed in block from I/O bus and metadata is accessed in byte-addressable manner from memory bus. Metadata access is significantly accelerated and metadata I/O is eliminated because metadata in NVM is not flushed back to disk periodically anymore. A light-weight consistency mechanism combining fine-grained versioning and transaction is introduced in the FSMAC. The FSMAC is implemented on the basis of Linux Ext4 file system and intensively evaluated under different workloads. Evaluation results show that the FSMAC accelerates file system up to 49.2 times for synchronized I/O and 7.22 times for asynchronized I/O. Jianxi Chen, Qingsong Wei, Cheng Chen 0008, Lingkun Wu |
MSST | 2 |
| 2013 | Improving flash-based disk cache with Lazy Adaptive ReplacementabstractThe increasing popularity of flash memory has changed storage systems. Flash-based solid state drive(SSD) is now widely deployed as cache for magnetic hard disk drives(HDD) to speed up data intensive applications. However, existing cache algorithms focus exclusively on performance improvements and ignore the write endurance of SSD. In this paper, we proposed a novel cache management algorithm for flash-based disk cache, named Lazy Adaptive Replacement Cache(LARC). LARC can filter out seldom accessed blocks and prevent them from entering cache. This avoids cache pollution and keeps popular blocks in cache for a longer period of time, leading to higher hit rate. Meanwhile, LARC reduces the amount of cache replacements thus incurs less write traffics to SSD, especially for read dominant workloads. In this way, LARC improves performance and extends SSD lifetime at the same time. LARC is self-tuning and low overhead. It has been extensively evaluated by both trace-driven simulations and a prototype implementation in flashcache. Our experiments show that LARC outperforms state-of-art algorithms and reduces write traffics to SSD by up to 94.5% for read dominant workloads, 11.2-40.8% for write dominant workloads. Sai Huang, Qingsong Wei, Jianxi Chen, Cheng Chen 0008, Dan Feng 0001 |
MSST | 2 |
| 2012 | HerpRap: A Hybrid Array Architecture Providing Any Point-in-Time Data Tracking for DatacenterabstractBoth physical disk failure and logical errors such as software error, user abuse and virus attacks may cause data lose. The risk of logical errors is far greater than physical disk failure. Moreover, existing RAID solution cannot satisfy the reliability requirement in face of the logical errors in data centers. It is therefore becoming increasingly important for RAID-based storage systems to be able to recover data to any point-in-time when logical errors occur. We proposed a novel storage array architecture, Herp Rap, which is able to recover data from both physical disk failure and logical errors. We have implemented a prototype of Herp Rap and carried out extensive performance measurements using DBT-2 and file system benchmarks. Our experiments demonstrated that the proposed Herp Rap is able to track or recover data to any point-in-time quickly by tracing back the history of block logs. Moreover, Herp Rap outperforms existing HDD-based or SSD-based RAID5 with copy-on-write (COW) snapshot in terms of performance, energy efficiency, failure recovery ability and reliability. Lingfang Zeng, Dan Feng 0001, Bo Mao 0003, Jianxi Chen, Qingsong Wei, Wenguo Liu 0004 |
CLUSTER | 5 |
| 2012 | HRAID6ML: A hybrid RAID6 storage architecture with mirrored loggingabstractThe RAID6 provides high reliability using double-parity-update at cost of high write penalty. In this paper, we propose HRAID6ML, a new logging architecture for RAID6 systems for enhanced energy efficiency, performance and reliability. HRAID6ML explores a group of Solid State Drives (SSDs) and Hard Disk Drives (HDDs): Two HDDs (parity disks) and several SSDs form RAID6. The free space of the two parity disks is used as mirrored log region of the whole system to absorb writes. The mirrored logging policy helps to recover system from parity disk failure. Mirrored logging operation does not introduce noticeable performance overhead to the whole system. HRAID6ML eliminates the additional hardware and energy costs, potential single point of failure and performance bottleneck. Furthermore, HRAID6ML prolongs the lifecycle of the SSDs and improves the systems energy efficiency by reducing the SSDs write frequency. We have implemented proposed HRAID6ML. Extensive trace-driven evaluations demonstrate the advantages of the HRAID6ML system over both traditional SSD-based RAID6 system and HDD-based RAID6 system. Lingfang Zeng, Dan Feng 0001, Jianxi Chen, Qingsong Wei, Bharadwaj Veeravalli, Wenguo Liu 0004 |
MSST | 4 |
| 2012 | SeWDReSS: on the design of an application independent, secure, wide-area disaster recovery storage system
Lingfang Zeng, Bharadwaj Veeravalli, Qingsong Wei, Dan Feng 0001 |
Multim. Tools Appl. | 3 |
| 2011 | WAFTL: A workload adaptive flash translation layer with data partitionabstractCurrent FTL schemes have inevitable limitations in terms of memory requirement, performance, garbage collection overhead, and scalability. To overcome these limitations, we propose a workload adaptive flash translation layer referred to as WAFTL. WAFTL explores either page-level or block-level address mapping for normal data block based on access patterns. Page Mapping Block (PMB) is used to store random data and handle large number of partial updates. Block Mapping Block (BMB) is utilized to store sequential data and lower overall mapping table. PMB or BMB is allocated on demand and the number of PMB or BMB eventually depends on workload. An efficient address mapping is designed to reduce overall mapping table and quickly conduct address translation. WAFTL explores a small part of flash space as Buffer Zone to log writes sequentially and migrate data into BMB or PMB based on threshold. Static and dynamic threshold setting are proposed to balance performance and mapping table size. WAFTL has been extensively evaluated under various enterprise workloads. Benchmark results conclusively demonstrate that proposed WAFTL is workload adaptive and achieves up to 80% performance improvement, 83% garbage collection overhead reduction and 50% mapping table reduction compared to existing FTL schemes. Qingsong Wei, Bozhao Gong, Suraj Pathak, Bharadwaj Veeravalli, Lingfang Zeng, Kanzo Okada |
MSST | 1 |
| 2010 | CDRM: A Cost-Effective Dynamic Replication Management Scheme for Cloud Storage ClusterabstractData replication has been widely used as a mean of increasing the data availability of large-scale cloud storage systems where failures are normal. Aiming to provide cost-effective availability, and improve performance and load-balancing of cloud storage, this paper presents a cost-effective dynamic replication management scheme referred to as CDRM. A novel model is proposed to capture the relationship between availability and replica number. CDRM leverages this model to calculate and maintain minimal replica number for a given availability requirement. Replica placement is based on capacity and blocking probability of data nodes. By adjusting replica number and location according to workload changing and node capacity, CDRM can dynamically redistribute workloads among data nodes in the heterogeneous cloud. We implemented CDRM in Hadoop Distributed File System (HDFS) and experiment results conclusively demonstrate that our CDRM is cost effective and outperforms default replication management of HDFS in terms of performance and load balancing for large-scale cloud storage. Qingsong Wei, Bharadwaj Veeravalli, Bozhao Gong, Lingfang Zeng, Dan Feng 0001 |
CLUSTER | 1 |
| 2010 | FlashCoop: A Locality-Aware Cooperative Buffer Management for SSD-Based Storage ClusterabstractRandom writes significantly limit the application of flash-based Solid State Drive (SSD) in enterprise environment due to its poor latency, negative impact on SSD lifetime and high garbage collection overhead. To release above limitations, we propose a locality-aware cooperative buffer scheme referred to as FlashCoop (Flash Cooperation), which leverages free memory of neighboring storage server to buffer writes over high speed network. Both temporal and sequential localities of access pattern are exploited in the design of cooperative buffer management. Leveraging the filtering effect of the cooperative buffer, FlashCoop can efficiently shape the I/O request stream and improve the sequentiality of the write accesses passed to the SSD. FlashCoop has been extensively evaluated under various enterprise workloads. Our benchmark results conclusively demonstrate that FlashCoop can achieve 52.3% performance improvement and 56.5% garbage collection overhead reduction compared to the system without FlashCoop. Qingsong Wei, Bozhao Gong, Suraj Pathak, Y. C. Tay |
ICPP | 1 |
| 2010 | Dynamic Replication Management for Object-Based Storage SystemabstractData replication has been widely used as a mean of increasing the data availability of large-scale storage systems where failures are normal. Aiming to provide cost-effective availability, and improve performance and load-balancing of large-scale storage cluster, this paper presents a dynamic replication management scheme referred to as DRM. A model is developed to express availability as function of replica number. Based on this model, minimal replica number to satisfy availability requirement can be determined. DRM further places these replicas among Object-Based Storage Devices (OSD) in a balance way, taking into account different capacity and blocking probability of each OSD in heterogeneous environment. Proposed DRM can dynamically redistribute workloads among OSD cluster by adjusting replica number and location according to workload changing and OSD capacity. Our experiment results conclusively demonstrate that DRM is reliable and can achieve a significant average response time, and load balancing for large-scale OSD cluster. Qingsong Wei, Bharadwaj Veeravalli |
NAS | 1 |
| 2009 | ARRAY: A Non-application-Related, Secure, Wide-Area Disaster Recovery Storage SystemabstractWith our society more information-driven, we have begun to distribute data in wide-area storage systems. At the same time, both physical failure and logic error have made it difficult to bring the necessary recovery to bear on remote data disaster, and understanding this proceeding. We describe ARRAY, a system architecture for data disaster recovery that combines reliability, storage space, and security to improve performance for data recovery applications. The paper presents an exhaustive analysis of the design space of ARRAY systems, focusing on the trade-offs between reliability, storage space, security, and performance that ARRAY must make. We present RSRAII (Replication-based Snapshot Redundant Array of Independent Imagefiles) which is a configurable RAID-like data erasure-coding, and also others benefits come from consolidation both erasure-coding and replication strategies. A novel algorithm is proposed to improve snapshot performance referred to as SMPDP (Snapshot based on Multi-Parallel Degree Pipeline). Lingfang Zeng, Dan Feng 0001, Bharadwaj Veeravalli, Qingsong Wei |
ISPA | 4 |
| 2009 | CPM: Cooperative power management for object-based storage clusterabstractDisk idle periods in server workload are short, which significantly limits the effectiveness of underline disk power management. To release this limitation, we present a cooperative power management (referred to as CPM) scheme to save energy with performance guarantee for object-based storage cluster. CPM reclaims idle memories of neighboring object-based storage devices (OSDs) over high speed network as remote cache to store evicted objects. Then requests missed in local cache could be hit by remote cache, and local disk does not necessarily spin back up to service these requests. Hence, CPM can artificially create long idle periods to provide more opportunities for underlying disk power management. CPM minimizes the risk of performance and energy penalty by spinning down disks only when predicted idle period is long enough to justify state-transition energy. Our rigorous experiment results conclusively demonstrate that CPM can dynamically adapt to workload changes and outperform existing solutions in terms of energy saving and performance for large-scale OSD cluster. Qingsong Wei, Bharadwaj Veeravalli |
MASCOTS | 1 |
| 2009 | Design and Performance Evaluation of a Versatile Object-Based File SystemabstractThe object-based storage system stripes data across large numbers of object-based storage devices (OSDs) to enable parallel data access. Subsequently, the workload presented to the individual OSD will be quite different from that of general purpose file systems. However, many distributed file systems employ general-purpose file systems as their underlying file system. This paper presents a versatile object-based file system (referred to as V-OBFS), an extent and B+tree based file system designed for use in OSDs. The V-OBFS implements flexible disk layout with multiple block size and differentiated free space management for both small objects and large objects. In addition, proposed V-OBFS enables fast object search and mapping between object ID and physical location with an efficient object namespace management. Our experiments show that our user-level implementation of V-OBFS can efficiently prevent file system fragmentation and outperforms Linux Ext2 and Ext3 by a factor of two or three. Qingsong Wei, Rajesh Vellore Arumugam, Kyawt Kyawt Khaing |
NPC | 1 |
| 2008 | DifferStore: A differentiated storage service in object-based storage systemabstractThis paper presents a differentiated storage service in object-based storage system, called DifferStore. To enable differentiated storage service for different applications in a single object-based storage platform, DifferStore utilizes a two-layer architecture to efficiently decouple upper-layer application specific storage policies and lower-layer application independent storage functions. For the lower application independent layer, this paper proposes a weight-based object I/O scheduler with differentiated scheduling policy for different request classes, and a versatile storage manager. The versatile storage manager implements differentiated storage policies in terms of disk layout and free space allocation, as well as an efficient object namespace management enabling directly access object on-disk data just with object ID. The DifferStore also provides ability for upper application specific layer to assign complex striping, placement, load-balancing policies and specific metadata structure of file. Experimental evaluation on our user space prototype demonstrates that the DifferStore can perform well under mixed workloads and satisfy requirements of different applications. Qingsong Wei |
CLUSTER | 1 |
| 2008 | DWC2: A dynamic weight-based cooperative caching scheme for object-based storage clusterabstractObject-based storage is emerging as a next generation of distributed storage technology. Aiming at improving the performance and load-balancing of large-scale object-based storage system, we present a dynamic weight-based cooperative caching scheme referred to as DWC2, which allows an object-based storage device (OSD) to use the available free cache of the neighbouring OSD. Our proposed DWC2replaces objects based on their weights which is a function of object size, popularity and replica number, and dynamically partitions the memory of OSD into local cache and remote cache according to activity workload. An object data is cached in local cache or remote cache of the cooperative OSDs, thus increasing cache hit ratio, reducing expensive disk access time as well as improving load balance. We benchmarked our proposed DWC2with existing cooperative caching schemes under various OSD environments. Our rigorous experiment results conclusively demonstrate that our DWC2is scalable and can achieve a significant cache hit ratio, average response time, and load balancing for large-scale OSD cluster. Qingsong Wei, Bharadwaj Veeravalli, Lingfang Zeng |
CLUSTER | 1 |
| 2007 | A high-speed and low-cost storage architecture based on virtual interface
Lingfang Zeng, Dan Feng 0001, Zhan Shi 0001, Jianxi Chen, Qingsong Wei |
Frontiers Comput. Sci. China | 5 |