Zuoru Yang

dblp:223/6886 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0001-6915-7100ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Encrypted Data Reduction: Removing Redundancy from Encrypted Data in Outsourced Storage
abstract
Storage savings and data confidentiality are two primary goals for outsourced storage. However, encryption by design destroys the content redundancy within plaintext data, so there exist design tensions when combining encryption with data reduction techniques (i.e., deduplication, delta compression, and local compression). We present EDRStore, an outsourced storage system that realizes encrypted data reduction to achieve both storage savings and data confidentiality. EDRStore’s core idea is a careful design of the encryption and data reduction workflows. It proposes new key generation and encryption schemes to preserve the content similarity of encrypted data for deduplication and delta compression. It further proposes selective local compression based on content similarity, so as to achieve storage savings of encrypted data from both delta compression and local compression. Evaluation on real-world datasets shows that EDRStore achieves higher storage savings than existing encrypted storage approaches and incurs moderate performance overhead compared with plaintext storage.
Zuoru Yang, Jingwei Li 0001, Patrick P. C. Lee
ACM Trans. Storage2
2022 Secure and Lightweight Deduplicated Storage via Shielded Deduplication-Before-Encryption
Zuoru Yang, Jingwei Li 0001, Patrick P. C. Lee
USENIX ATC1
2022 Enabling Secure and Space-Efficient Metadata Management in Encrypted Deduplication
abstract
Encrypted deduplication combines encryption and deduplication in a seamless way to provide confidentiality guarantees for the physical data in deduplicated storage, yet it incurs substantial metadata storage overhead due to the additional storage of keys. We present a new encrypted deduplication storage system called${\sf Metadedup}$, which suppresses metadata storage by also applying deduplication to metadata. Its idea builds on indirection, which adds another level of metadata chunks that record metadata information. We find that metadata chunks are highly redundant in real-world workloads and hence can be effectively deduplicated. We further extend${\sf Metadedup}$to incorporate multiple servers via a distributed key management approach, so as to provide both fault-tolerant storage and security guarantees. We extensively evaluate${\sf Metadedup}$from performance and storage efficiency perspectives. We show that${\sf Metadedup}$achieves high throughput in writing and restoring files, and saves the metadata storage by up to 93.94 percent for real-world backup workloads.
Jingwei Li 0001, Suyu Huang, Yanjing Ren, Zuoru Yang, Patrick P. C. Lee, Xiaosong Zhang 0001, Yao Hao
IEEE Trans. Computers4
2022 Tunable Encrypted Deduplication with Attack-resilient Key Management
abstract
Conventional encrypted deduplication approaches retain the deduplication capability on duplicate chunks after encryption by always deriving the key for encryption/decryption from the chunk content, but such a deterministic nature causes information leakage due to frequency analysis. We present TED , a tunable encrypted deduplication primitive that provides a tunable mechanism for balancing the tradeoff between storage efficiency and data confidentiality. The core idea of TED is that its key derivation is based on not only the chunk content but also the number of duplicate chunk copies, such that duplicate chunks are encrypted by distinct keys in a controlled manner. In particular, TED allows users to configure a storage blowup factor, under which the information leakage quantified by an information-theoretic measure is minimized for any input workload. In addition, we extend TED with a distributed key management architecture and propose two attack-resilient key generation schemes that trade between performance and fault tolerance. We implement an encrypted deduplication prototype TEDStore to realize TED in networked environments. Evaluation on real-world file system snapshots shows that TED effectively balances the tradeoff between storage efficiency and data confidentiality, with small performance overhead.
Zuoru Yang, Jingwei Li 0001, Yanjing Ren, Patrick P. C. Lee
ACM Trans. Storage1
2021 Accelerating Encrypted Deduplication via SGX
Yanjing Ren, Jingwei Li 0001, Zuoru Yang, Patrick P. C. Lee, Xiaosong Zhang 0001
USENIX ATC3
2021 Network-Wide Forwarding Anomaly Detection and Localization in Software Defined Networks
abstract
A crucial requirement for Software Defined Network (SDN) is that data plane forwarding behaviors should always agree with control plane policies. Such requirement cannot be met when there areforwarding anomalies, where packets deviate from the paths specified by the controller. Most anomaly detection methods for SDN install dedicated rules to collect statistics of each flow, and check whether the statistics conform to the “flow conservation principle”. We find these methods have a limited detection scope: they look at one flow each time, thus can only check a small number of flows simultaneously. In addition, dedicated rules for statistics collection can impose a large overhead on flow tables of SDN switches. To this end, this paper presents FOCES, a network-wide forwarding anomaly detection and localization method in SDN. Different from previous methods, FOCES applies a new kind of flow conservation principle at network wide, and can check forwarding behaviors ofallflows in the network simultaneously, without installing any dedicated rules. Finally, FOCES applies a voting-based method to localize malicious switches when anomalies are detected. Experiments with four network topologies show that FOCES can achieve a detection precision higher than 90%, when the packet loss rate is no larger than 10%, and a localization accuracy of around 80% when the packet loss rate is no larger than 5%.
Peng Zhang 0011, Fangzheng Zhang, Shimin Xu, Zuoru Yang, Hao Li 0011, Qi Li 0002, Huanzhao Wang, Chao Shen 0001, Chengchen Hu
IEEE/ACM Trans. Netw.4
2021 Repair Pipelining for Erasure-coded Storage: Algorithms and Evaluation
abstract
We propose repair pipelining , a technique that speeds up the repair performance in general erasure-coded storage. By carefully scheduling the repair of failed data in small-size units across storage nodes in a pipelined manner, repair pipelining reduces the single-block repair time to approximately the same as the normal read time for a single block in homogeneous environments. We further design different extensions of repair pipelining algorithms for heterogeneous environments and multi-block repair operations. We implement a repair pipelining prototype, called ECPipe , and integrate it as a middleware system into two versions of Hadoop Distributed File System (HDFS) (namely, HDFS-RAID and HDFS-3) as well as Quantcast File System. Experiments on a local testbed and Amazon EC2 show that repair pipelining significantly improves the performance of degraded reads and full-node recovery over existing repair techniques.
Xiaolu Li 0002, Zuoru Yang, Runhui Li, Patrick P. C. Lee, Qun Huang 0001, Yuchong Hu
ACM Trans. Storage2
2020 Balancing storage efficiency and data confidentiality with tunable encrypted deduplication
abstract
Conventional encrypted deduplication approaches retain the deduplication capability on duplicate chunks after encryption by always deriving the key for encryption/decryption from the chunk content, but such a deterministic nature causes information leakage due to frequency analysis. We present TED, a tunable encrypted deduplication primitive that provides a tunable mechanism for balancing the tradeoff between storage efficiency and data confidentiality. The core idea of TED is that its key derivation is based on not only the chunk content but also the number of duplicate chunk copies, such that duplicate chunks are encrypted by distinct keys in a controlled manner. In particular, TED allows users to configure a storage blowup factor, under which the information leakage quantified by an information-theoretic measure is minimized for any input workload. We implement an encrypted deduplication prototype TEDStore to realize TED in networked environments. Evaluation on real-world file system snapshots shows that TED effectively balances the trade-off between storage efficiency and data confidentiality, with small performance overhead.
Jingwei Li 0001, Zuoru Yang, Yanjing Ren, Patrick P. C. Lee, Xiaosong Zhang 0001
EuroSys2
2018 FOCES: Detecting Forwarding Anomalies in Software Defined Networks
abstract
A crucial requirement for Software Defined Network (SDN) is that data plane forwarding behaviors should always agree with control plane policies. Such requirement cannot be met when there are forwarding anomalies, where packets deviate from the paths specified by the controller. Most anomaly detection methods for SDN install dedicated rules to collect statistics of each flow, and check whether the statistics conform to the flow conservation principle. Such per-flow detection methods have a limited detection scope: they look at one flow each time, thus can only check a limited number of flows simultaneously. In addition, dedicated rules for statistics collection can impose a large overhead on flow tables of SDN switches. To this end, this paper presents FOCES, a network-wide forwarding anomaly detection method in SDN. Different from previous methods, FOCES applies a new kind of flow conservation principle at network wide, and can check forwarding behaviors of all flows in the network simultaneously, without installing any dedicated rules. Experiments show FOCES can achieve a detection precision higher than 90% for four network topologies, even when packet loss rates are as high as 10%.
Peng Zhang 0011, Shimin Xu, Zuoru Yang, Hao Li 0011, Qi Li 0002, Huanzhao Wang, Chengchen Hu
ICDCS3