VLDB 2026 Research / reviewers in the wild / expert
Yinjin Fu
dblp:93/6108
· DBLP profile ↗
28ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0001-9107-1338ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 first-author · 7 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedGraph-ID: A Federated Graph Learning Framework for Intrusion Detection in UAV Networks Under Adversarial Settings
Qingli Zeng, Yinjin Fu, Farid Naït-Abdesselam |
INFOCOM | 2 |
| 2025 | SDXE: Accelerating Secure Cloud Deduplication Via SGX in Edge ComputingabstractSecure data deduplication in edge computing has emerged as a pivotal technique to enhance the cost-efficiency and data security of cloud storage infrastructures. However, the prevailing methodologies predominantly rely on resource-intensive cryptographic operations, resulting in substantial communication and computation overheads. In this paper, we introduce SDXE: a secure data deduplication framework that leverages cloud-edge collaborative computing through Intel's Software Guard Extensions (SGX). It can strike a balance between system performance and data security by harnessing SGX to protect sensitive operations, circumventing the need for conventional cryptographic algorithms. Additionally, we propose a hierarchical storage strategy predicated on data temperature, alongside a version-based two-level fingerprint index structure, to optimize data storage efficiency and enhance transfer performance. Comparing with the typical cryptography-based cloud-edge collaborative secure deduplication schemes, our experimental results demonstrate that SDXE can significantly enhance data communication efficiency with high data security, achieving a remarkable$9.25 \times$upload throughput in client-edge stage,$3.44 \times$upload throughput in edge-cloud stage and$2.96 \times$download throughput. Yinjin Fu, Nong Xiao 0001 |
ICC | 2 |
| 2025 | RT-PMalloc: Optimizing Persistent Memory Allocation for Soft Real-Time SystemsabstractPredictable dynamic allocation of Persistent Memory (PM) is pivotal to enable flexible and maintainable software design in PM-aware real-time systems. Existing PM allocators suffer from significant variability in the response times of allocation/deallocation operations. Furthermore, although there exist several real-time allocators for DRAM, designers of PM allocators have to tackle the additional challenging problem of crash consistency in order to fully utilize the non-volatile nature of PM. In this paper, we present RT-PMalloc, a novel persistent memory allocator designed for soft real-time applications in server environments where bounded allocation latency is critical. RT-PMalloc builds upon Makalu and introduces three targeted optimizations-profile-guided pre-allocation, constant-time header indexing, and refactored block clearing-to improve response time predictability while preserving crash consistency. Our evaluation shows that RT-PMalloc reduces worst-case allocation latency by up to 85.65% compared to the state-of-the-art PM allocator and maintains tighter latency bounds under multi-threaded workloads. These results demonstrate the feasibility of predictable dynamic PM management for modern soft real-time systems. Yuquan Chi, Yinjin Fu, Nong Xiao 0001 |
ICCD | 2 |
| 2025 | RABBIT: Managing Hierarchical Memory with Intelligent Tiering Aware DeduplicationabstractMemory tiering is a solution to classify the increasing data for real-time analysis in hierarchical storage management based on its usage patterns and importance.However, traditional storage tiering methods that utilize static thresholds or heuristics are rigid and inefficient, while most existing AI methods use overly complex prediction models, which cost a lot of resources in training and classifying.To address these issues, we propose RABBIT, a Rarely Accessed Blocks Based Intelligent Tiering method for tiered-aware deduplication memory systems.It introduces a monitoring process that can track file-level access features in real systems, predicts which memory layer each data blocks should enter with a daily updated decision tree model, and dynamically migrates the chosen data in a corresponding post-process with error correction mechanism.Meanwhile, a block-level global deduplication scheme is used to ensure that only unique data blocks occupy space in the entire tiered memory system and save storage resources.We build a three-layers and four-layers architecture to evaluate our RABBIT design with block-level I/O traces and real-world workloads and prove RABBIT to be a highly scalable method.Compared with stateof-the-art storage tiering methods, our experimental results show that RABBIT can reduce the average per block latency by 16.89% to 21.84% with a higher access hit rate and save storage capacity by 8.7% to 21.4% for learning access features at the block-level and file-level, respectively. Zilu Yao, Yinjin Fu, Nong Xiao 0001 |
Internetware | 2 |
| 2024 | Combining Buffered I/O and Direct I/O in Distributed File Systems
Yingjin Qian, Marc-Andre Vef, Patrick Farrell, Andreas Dilger, Shuichi Ihara, Yinjin Fu, André Brinkmann |
FAST | 7 |
| 2024 | Multi-Stage Dynamic Cuckoo FiltersabstractDeduplication is a highly efficient data reduction technique to improve storage efficiency and save costs. However, the deduplication performance is severely affected by the limited main memory capacity and the disk access bottleneck for the increasing chunk index lookup. In this paper, we propose Multi-Stage Dynamic Cuckoo Filters (MDCF) to speed up index lookup in deduplication systems. Firstly, MDCF modifies the deletion algorithm of standard cuckoo filter to solve the potential problem that it could not identify whether a cell is empty or not. Secondly, MDCF adjusts the insertion algorithm and expands capacity for the increasing load at bucket level or cuckoo filter level. What is more, MDCF integrates a BitSet and CountingSet for each bucket, which significantly reduces the memory overhead and improves the search efficiency without affecting the false positive rate. The experimental results show that compared with DCF, the query, insertion and deletion efficiency of MDCF are improved by about 55.5%, 20.0% and 21.4% respectively. While MDCF's false positive rate was 49.4% lower than that of DBF. And the data backup throughput can be increased by about 163.0% with MDCF to accelerate index lookup. Yinjin Fu, Nong Xiao 0001 |
ICCD | 2 |
| 2024 | Unified Lossless-Throughput Architecture for AES and SM4 Encryption with Changeable KeysabstractNetwork devices targeting to implement data-intensive applications often require the outstanding performance of symmetric encryption, when dealing with multiple concurrent requests from multiple users. Despite the numerous works on high-performance implementation of AES and SM4, hybrid architectures with lossless throughput when the key changes have not been proposed. In this paper, we propose a unified fully-pipelined architecture of AES and SM4 targeting high-performance Galois/Counter Mode application scenarios. The architecture is able to maintain the consistent throughput of input and output datastreams with changeable keys. Compared with state-of-the-art works implemented with the TSMC 65nm process, our design can reduce the area by 26.83% by using a shared composite S-box. With the one-hot S-box, our design can reduce power consumption by 30.93% and increase throughput by 34.21%. Zhishuo Huang, Haosong Zhao, Donald Donglong Chen, Shuyan Zhu, Yinjin Fu, Nong Xiao 0001, Yao Liu 0006 |
ISCAS | 6 |
| 2023 | Community Detection-Empowered Hybrid Network Slicing for Aerial Communication ServicesabstractNetwork slicing is the most promising paradigm that enables the provision of services for aerial communication tasks, such as data traffic offloading, disaster relief, and data collection, with one shared physical network. However, current researchers primarily focus on the modeling of utilizing service function chaining (SFC) or task offloading, which have inherent limitations for areial communication tasks. While the former restricts the source-sink nodes to be fixed, the latter models network slicing as a whole. To further unleash the flexibility of network slicing, we introduce the hybrid network slicing (HNS) modeling in this paper. In our HNS, each aerial service is abstracted as multiple SFCs with dispersed source-sink nodes. These SFCs can be flexibly deployed in a unified or split manner based on the service workload and Quality of Service (QoS) requirements. Then, we formulate the HNS problem in a multi-tier network system as a mixed integer programming (MILP) problem and propose exact and community detection-based heuristic approaches to address the proposed problem. The simulation results demonstrate that the proposed heuristic approaches can effectively mitigate computational complexity than the exact method while providing near-optimal solutions for aerial communication tasks. Chenjing Tian, Haotong Cao, Yinjin Fu, Xijian Luo |
GLOBECOM | 4 |
| 2023 | GreDedup: A Greedy-Based Application-Aware Data Routing Strategy for Distributed DeduplicationabstractWe propose GreDedup, a greedy algorithm based application-aware data routing strategy for distributed deduplication, which can achieve a good tradeoff between high global deduplication ratio and scalable performance by reducing the communication overhead and avoiding disk bottleneck. We extract semantic information to classify backup files, and use the greedy algorithm to route files with the same type to as few storage servers as possible with the help of application tables. In intra-node deduplication, we maintain a unique chunk fingerprint index for each file type to reduce disk access times. We perform experiments to compare GreDedup with state-of-the-art alternatives under public datasets. The results show that GreDedup can achieve high global deduplication ratio almost the same as the high overhead scheme, but its write performance even exceeds that of the low overhead method with good load balancing. Yinjin Fu, Nong Xiao 0001, Yingjin Qian |
ICPADS | 2 |
| 2023 | Online and reliable SFC protection scheme of distributed cloud network for future IoT application
Chenjing Tian, Haotong Cao, Yinjin Fu, Sahil Garg, Georges Kaddoum, Mohammad Mehedi Hassan |
Comput. Commun. | 3 |
| 2022 | Characterizing and Optimizing Hybrid DRAM-PM Main Memory System with Application AwarenessabstractPersistent memory (PM) has always been used in combination with DRAM to configure hybrid main memory systems that can obtain both the high performance of DRAM and large capacity of PM. There are critical management challenges in data placement, memory concurrency and workload scheduling for the concurrent execution of multiple application workloads. But the non-negligible performance gap between DRAM and PM makes the existing application-agnostic management strategies inefficient in reaching the full potential of hybrid memory. In this paper, we propose a series of application aware optimization strategies, including application aware data placement, adaptive thread allocation and inter-application interference avoiding, to improve the concurrent performance of different application workloads on hybrid memory. Finally, we provide the performance evaluation for our application aware solutions on real hybrid memory hardware with some comprehensive benchmark suites. Our experimental results show that the duration of multi-application concurrent execution on hybrid memory can be reduced by at most 60.7% for application aware data placement, 37.7% for adaptive thread allocation and 34.8% for workload scheduling with inter-application interference avoiding, respectively. And the additive effects of all these three optimization methods can reach 62.8% performance improvement with negligible overheads. Yongfeng Wang, Yinjin Fu, Zhiguang Chen 0001, Nong Xiao 0001 |
DATE | 2 |
| 2022 | Fog-to-MultiCloud Cooperative Ehealth Data Management With Application-Aware Secure DeduplicationabstractThe healthcare industry faces challenges regarding the security and efficiency of data management for patient health records (PHRs). We propose SafePHR: a secure and efficient medical data service with application aware deduplication in fog-to-multicloud encrypted storage. It introduces a fog-to-multicloud cooperative storage model, which combines low-latency and high-safety local fog with unlimited capacity and built-in disaster recovery in remote multi-cloud to enhance eHealth data management. It also provides an application aware secure deduplication scheme to improve the data traffic and space efficiency of cloud storage for encrypted PHRs using variants of convergent encryption. Then, it further enhances the resiliency and security of cloud storage by striping encrypted data across multiple cloud vendors with a fault-tolerant coding scheme. Compared with the state-of-the-art of cloud-assisted eHealth systems, our experiments demonstrate that SafePHR can ensure data confidentiality of deduplication, achieve a competitive low cloud storage space overhead, and improve system performance with high fault-tolerance capability. Yinjin Fu, Nong Xiao 0001, Tao Chen 0005, Jian Wang 0025 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Design and Simulation of Content-Aware Hybrid DRAM-PCM Memory SystemabstractPhase Change Memory (PCM) can directly connect persistent memory to main memory bus, while it achieves high read throughput and low standby power, the critical concerns are its poor write performance and limited durability. A naturally in-spired design is the hybrid memory architecture that fuses DRAM and PCM, so as to exploit the positive aspects of both types of memory. Unfortunately, existing solutions are seriously challenged by the limited main memory size, which is the primary bottleneck of in-memory computing. In this paper, we introduce a novel Content Aware Hybrid DRAM-PCM memory system framework—CAHRAM, which exploits deduplication to improve line sharing with high memory efficiency. It reduces write traffic to hybrid memory by removing unnecessary duplicate line writes, thereby further enhancing the write endurance of PCM. And it also substantially extends available free memory space by coalescing redundant lines in hybrid memory. We also design a reference-based page migration technique to minimize the access overheads caused by the performance gap between DRAM and PCM. Compared with the state-of-the-art in a hybrid memory simulator, our experiment results show that CAHRAM can achieve the highest I/O performance and the longest PCM lifetime with the competitive efficiencies in space and energy. Yinjin Fu, Yutong Lu, Zhiguang Chen 0001, Nong Xiao 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | CARAM: A Content-Aware Hybrid PCM/DRAM Main Memory System Framework
Yinjin Fu |
NPC | 1 |
| 2019 | F2MC: Enhancing Data Storage Services with Fog-toMultiCloud Hybrid ComputingabstractPublic cloud storage can provide customers with unlimited capacity and built-in disaster recovery by exploiting IT resources in data center but suffers high latency and security risk. The emergency of fog computing can mitigate data security risk by leveraging local private server with low latency. We introduce F2MC: a Fog-to-MultiCloud hybrid storage service that combines local fog computing with remote cloud computing to enhance the quality of service (QoS) of data management. It provides an application aware data reduction scheme to improve the data traffic and space efficiency of cloud storage by exploiting deduplication and compression. Then, it further enhances the resiliency and security of cloud storage by striping encrypted data across multiple cloud vendors with a secure fault-tolerant coding. Compared with the state-of-the-art of multicloud storage systems, our experiments demonstrate that F2MC can achieve the highest performance with the same fault-tolerance capability, and a competitive low cloud cost due to its low cloud storage space overheads. Yinjin Fu, Xiaofeng Qiu, Jian Wang 0025 |
IPCCC | 1 |
| 2019 | Application-Aware Big Data Deduplication in Cloud EnvironmentabstractDeduplication has become a widely deployed technology in cloud data centers to improve IT resources efficiency. However, traditional techniques face a great challenge in big data deduplication to strike a sensible tradeoff between the conflicting goals of scalable deduplication throughput and high duplicate elimination ratio. We propose AppDedupe, an application-aware scalable inline distributed deduplication framework in cloud environment, to meet this challenge by exploiting application awareness, data similarity and locality to optimize distributed deduplication with inter-node two-tiered data routing and intra-node application-aware deduplication. It first dispenses application data at file level with an application-aware routing to keep application locality, then assigns similar application data to the same storage node at the super-chunk granularity using a handprinting-based stateful data routing scheme to maintain high global deduplication efficiency, meanwhile balances the workload across nodes. AppDedupe builds application-aware similarity indices with super-chunk handprints to speedup the intra-node deduplication process with high efficiency. Our experimental evaluation of AppDedupe against state-of-the-art, driven by real-world datasets, demonstrates that AppDedupe achieves the highest global deduplication efficiency with a higher global deduplication effectiveness than the high-overhead and poorly scalable traditional scheme, but at an overhead only slightly higher than that of the scalable but low duplicate-elimination-ratio approaches. Yinjin Fu, Nong Xiao 0001, Hong Jiang 0001, Guyu Hu |
IEEE Trans. Cloud Comput. | 1 |
| 2018 | One Size Does Not Fit All: The Case for Chunking Configuration in Backup DeduplicationabstractData backup is regularly required by both enterprise and individual users to protect their data from unexpected loss. There are also various commercial data deduplication systems or software that help users to eliminate duplicates in their backup data to save storage space. In data deduplication systems, the data chunking process splits data into small chunks. Duplicate data is identified by comparing the fingerprints of the chunks. The chunk size setting has significant impact on deduplication performance. A variety of chunking algorithms have been proposed in recent studies. In practice, existing systems often set the chunking configuration in an empirical manner. A chunk size of 4KB or 8KB is regarded as the sweet spot for good deduplication performance. However, the data storage and access patterns of users vary and change along time, as a result, the empirical chunk size setting may not lead to a good deduplication ratio and sometimes results in difficulties of storage capacity planning. Moreover, it is difficult to make changes to the chunking settings once they are put into use as duplicates in data with different chunk size settings cannot be eliminated directly. In this paper, we propose a sampling-based chunking method and develop a tool named SmartChunker to estimate the optimal chunking configuration for deduplication systems. Our evaluations on real-world datasets demonstrate the efficacy and efficiency of SmartChunker. Huijun Wu 0001, Chen Wang 0008, Kai Lu 0001, Yinjin Fu, Liming Zhu 0001 |
CCGrid | 4 |
| 2018 | A Differentiated Caching Mechanism to Enable Primary Storage Deduplication in CloudsabstractExisting primary deduplication techniques either use inline caching to exploit locality in primary workloads or use post-processing deduplication to avoid the negative impact on I/O performance. However, neither of them works well in the cloud servers running multiple services for the following two reasons: First, the temporal locality of duplicate data writes varies among primary storage workloads, which makes it challenging to efficiently allocate the inline cache space and achieve a good deduplication ratio. Second, the post-processing deduplication does not eliminate duplicate I/O operations that write to the same logical block address as it is performed after duplicate blocks have been written. A hybrid deduplication mechanism is promising to deal with these problems. Inline fingerprint caching is essential to achieving efficient hybrid deduplication. In this paper, we present a detailed analysis of the limitations of using existing caching algorithms in primary deduplication in the cloud. We reveal that existing caching algorithms either perform poorly or incur significant memory overhead in fingerprint cache management. To address this, we propose a novel fingerprint caching mechanism that estimates the temporal locality of duplicates in different data streams and prioritizes the cache allocation based on the estimation. We integrate the caching mechanism and build a hybrid deduplication system. Our experimental results show that the proposed mechanism provides significant improvement for both deduplication ratio and overhead reduction. Huijun Wu 0001, Chen Wang 0008, Yinjin Fu, Sherif Sakr, Kai Lu 0001, Liming Zhu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | DualStack: A High Efficient Dynamic Page Scheduling Scheme in Hybrid Main MemoryabstractWith the development of big data and multi-core processors technology, DRAM-only based main memory cannot satisfy the requirements of in-memory computing in high memory capacity and low energy consumption. The emerging memory technology-phase change memory (PCM) is proposed to break the bottleneck of the current memory system. However, its weaknesses in write endurance and long access latency make it cannot fully replace DRAM. Consequently, researchers presented the architectural design aimed at DRAM/ PCM hybrids and the corresponding page migration scheme to give full play to their merits. The urgent challenges facing by existed page migration schemes are poor performed under weak locality in data streams and further improvement need in prediction of future access tendency. In this paper, we propose an efficient page migration policy called DualStack which features dynamic page management according to global read and write information and temporal locality. It is designed to keep write-intensive pages to DRAM and read-intensive pages to PCM, and specially avoid frequent and unnecessary migration between hybrid memory media. Compared to the state-of-the-art of hybrid main memory, our experimental results indicate that DualStack can effectively improve the system I/O latency by 38%~58% on the premise of reducing the system power consumption by 20%~30%. Yinjin Fu, Guyu Hu |
NAS | 2 |
| 2017 | Fast counting the cardinality of flows for big traffic over sliding windows
Jingsong Shan, Yinjin Fu, Guiqiang Ni, Jianxin Luo, Zhaofeng Wu |
Frontiers Comput. Sci. | 2 |
| 2015 | Tunneling-based Multi-path Routing Mechanism in Packet-Switched Non-Geostationary Satellite Networks
Guyu Hu, Zhaofeng Wu, Fenglin Jin, Bowei Yang, Yinjin Fu |
ICA3PP (4) | 6 |
| 2014 | Read-Performance Optimization for Deduplication-Based Storage Systems in the CloudabstractData deduplication has been demonstrated to be an effective technique in reducing the total data transferred over the network and the storage space in cloud backup, archiving, and primary storage systems, such as VM (virtual machine) platforms. However, the performance of restore operations from a deduplicated backup can be significantly lower than that without deduplication. The main reason lies in the fact that a file or block is split into multiple small data chunks that are often located in different disks after deduplication, which can cause a subsequent read operation to invoke many disk IOs involving multiple disks and thus degrade the read performance significantly. While this problem has been by and large ignored in the literature thus far, we argue that the time is ripe for us to pay significant attention to it in light of the emerging cloud storage applications and the increasing popularity of the VM platform in the cloud. This is because, in a cloud storage or VM environment, a simple read request on the client side may translate into a restore operation if the data to be read or a VM suspended by the user was previously deduplicated when written to the cloud or the VM storage server, a likely scenario considering the network bandwidth and storage capacity concerns in such an environment. To address this problem, in this article, we propose SAR, an SSD (solid-state drive)-Assisted Read scheme, that effectively exploits the high random-read performance properties of SSDs and the unique data-sharing characteristic of deduplication-based storage systems by storing in SSDs the unique data chunks with high reference count, small size, and nonsequential characteristics. In this way, many read requests to HDDs are replaced by read requests to SSDs, thus significantly improving the read performance of the deduplication-based storage systems in the cloud. The extensive trace-driven and VM restore evaluations on the prototype implementation of SAR show that SAR outperforms the traditional deduplication-based and flash-based cache schemes significantly, in terms of the average response times. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Yinjin Fu, Lei Tian 0001 |
ACM Trans. Storage | 4 |
| 2014 | Application-Aware Local-Global Source Deduplication for Cloud Backup Services of Personal StorageabstractIn personal computing devices that rely on a cloud storage environment for data backup, an imminent challenge facing source deduplication for cloud backup services is the low deduplication efficiency due to a combination of the resource-intensive nature of deduplication and the limited system resources. In this paper, we present ALG-Dedupe, an Application-aware Local-Global source deduplication scheme that improves data deduplication efficiency by exploiting application awareness, and further combines local and global duplicate detection to strike a good balance between cloud storage capacity saving and deduplication time reduction. We perform experiments via prototype implementation to demonstrate that our scheme can significantly improve deduplication efficiency over the state-of-the-art methods with low system overhead, resulting in shortened backup window, increased power efficiency and reduced cost for cloud backup services of personal storage. Yinjin Fu, Hong Jiang 0001, Nong Xiao 0001, Lei Tian 0001, Fang Liu 0002, Lei Xu 0038 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Application-Aware Client-Side Data Reduction and Encryption of Personal Data in Cloud Backup Services
Yinjin Fu, Nong Xiao 0001, Xiangke Liao, Fang Liu 0002 |
J. Comput. Sci. Technol. | 1 |
| 2012 | A Scalable Inline Cluster Deduplication Framework for Big Data Protection
Yinjin Fu, Hong Jiang 0001, Nong Xiao 0001 |
Middleware | 1 |
| 2012 | SAR: SSD Assisted Restore Optimization for Deduplication-Based Storage Systems in the CloudabstractThe explosive growth of digital content results in enormous strains on the storage systems in the cloud environment. The data deduplication technology has been demonstrated to be very effective in shortening the backup window and saving the network bandwidth and storage space in cloud backup, archiving and primary storage systems such as VM platforms. However, the delay and power consumption of the restore operations from a deduplicated storage can be significantly higher than those without deduplication. The main reason lies in the fact that a file or block is split into multiple small data chunks that are often located in non-sequential locations on HDDs after deduplication, which can cause a subsequent read operation to invoke many HDD I/O requests involving multiple disk seeks. To address this problem, in this paper we propose SAR, an SSD Assisted Restore scheme, that effectively exploits the high random-read performance and low power-consumption properties of SSDs and the unique data sharing characteristic of deduplication-based storage system by storing in SSDs the unique data chunks with high reference count, small size and non-sequential characteristics. In this way, many critical random-read requests to HDDs are replaced by read requests to SSDs, thus significantly improving the system performance and energy efficiency. The extensive trace-driven and VM restore evaluations on the prototype implementation of SAR show that SAR outperforms the traditional deduplication-based schemes significantly, in terms of both restore performance and energy efficiency. Bo Mao 0003, Hong Jiang 0001, Suzhen Wu, Yinjin Fu, Lei Tian 0001 |
NAS | 4 |
| 2011 | AA-Dedupe: An Application-Aware Source Deduplication Approach for Cloud Backup Services in the Personal Computing EnvironmentabstractThe market for cloud backup services in the personal computing environment is growing due to large volumes of valuable personal and corporate data being stored on desktops, laptops and smart phones. Source deduplication has become a mainstay of cloud backup that saves network bandwidth and reduces storage space. However, there are two challenges facing deduplication for cloud backup service clients: (1) low deduplication efficiency due to a combination of the resource-intensive nature of deduplication and the limited system resources on the PC-based client site, and (2) low data transfer efficiency since post-deduplication data transfers from source to backup servers are typically very small but must often cross a WAN. In this paper, we present AA-Dedupe, an application-aware source deduplication scheme, to significantly reduce the computational overhead, increase the deduplication throughput and improve the data transfer efficiency. The AA-Dedupe approach is motivated by our key observations of the substantial differences among applications in data redundancy and deduplication characteristics, and thus is based on an application-aware index structure that effectively exploits this application awareness. Our experimental evaluations, based on an AA-Dedupe prototype implementation, show that our scheme can improve deduplication efficiency over the state-of-art source-deduplication methods by a factor of 2-7, resulting in shortened backup window, increased power-efficiency and reduced cost for cloud backup services. Yinjin Fu, Hong Jiang 0001, Nong Xiao 0001, Lei Tian 0001, Fang Liu 0002 |
CLUSTER | 1 |
| 2008 | A Novel Dynamic Metadata Management Scheme for Large Distributed Storage SystemsabstractIn large distributed storage systems, metadata is usually managed separately by a metadata server cluster. The partitioning of the metadata among the servers is of critical importance for maintaining efficient MDS operation and a desirable load distribution across the cluster. We present a dynamic directory partitioning (DDP) metadata management scheme, directory metadata and file metadata are managed in different ways, and the dynamically changing workload can be balanced by adjusting the metadata distribution on metadata servers. Our simulation results show that our approach, comparing with other metadata management strategies, has advantages in performance, scalability and adaptability. Yinjin Fu, Nong Xiao 0001, Enqiang Zhou |
HPCC | 1 |