EDBT 2026 Demo / reviewers in the wild / expert
Yanjing Ren
dblp:252/6063
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-8882-6293ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SACK: Shielding Dynamic Attribute-based Access Control in Persistent Key-Value Stores
Yanjing Ren, Jingwei Li 0001, Patrick Lee |
Proc. VLDB Endow. | 1 |
| 2025 | SGX-Enabled Encrypted Cross-Cloud Data SynchronizationabstractThe increasing adoption of multicloud storage has necessitated the development of efficient cross-cloud data synchronization to improve performance and accessibility across regions, and reduce reliance on single cloud service. Yet, securing cross-cloud synchronization while achieving network efficiency poses challenges. First, ensuring data confidentiality and integrity is difficult due to the diverse encryption configurations and management complexities inherent to different clouds. Second, balancing security and network efficiency is non-trivial, as encryption disrupts content redundancy, complicating efforts to reduce network traffic. We present SeedSync, a system designed to provide secure and efficient cross-cloud data synchronization. SeedSync leverages shielded execution to ensure both security guarantees and network efficiency. It enables encrypted data synchronization without revealing sensitive information and ensures end-to-end data integrity with limited performance overhead using tree-structured integrity protection. It also addresses the dilemma between encryption and network reduction by allowing data to be processed unencrypted within a secure shielded region. Furthermore, inspired by workload characteristics, it speeds up data synchronization by reducing fine-grained network transmission and context switching of SGX. Evaluation on real-world datasets shows that SeedSync achieves up to 8.2 × higher throughput and 5.4 × network reduction existing traffic compared with encrypted synchronization approaches, while incurring limited overhead compared with plaintext synchronization. Yanjing Ren, Jingwei Li 0001, Patrick P. C. Lee |
ICDCS | 2 |
| 2024 | ELECT: Enabling Erasure Coding Tiering for LSM-tree-based Storage
Yanjing Ren, Yuanming Ren, Xiaolu Li 0002, Yuchong Hu, Jingwei Li 0001, Patrick P. C. Lee |
FAST | 1 |
| 2024 | Enhancing LSM-Tree Key-Value Stores for Read-Modify-Writes via Key-Delta SeparationabstractRead-modify-writes (RMWs) are increasingly observed in practical key-value (KV) storage workloads to support fine-grained updates. To make RMWs efficient, one approach is to write deltas (i.e., changes to current values) to the log-structured merge-tree (LSM-tree), yet it increases the read overhead caused by retrieving and combining a chain of deltas. We propose a notion called key-delta (KD) separation to support efficient reads and RMWs in LSM-tree KV stores under RMW-intensive workloads. KD separation aims to store deltas in separate storage areas and group deltas into storage units called buckets, such that all deltas of a key are kept in the same bucket and can be all together accessed in subsequent reads. To this end, we build KDSep, a middleware layer that realizes KD separation and integrates KDSep into state-of-the-art LSM-tree KV stores (e.g., RocksDB and BlobDB). We show that KDSep achieves significant I/O throughput gains and read latency reduction under RMW-intensive workloads while preserving the efficiency in general workloads. Yanjing Ren, Shujie Han 0001, Patrick P. C. Lee |
ICDE | 2 |
| 2024 | CDCache: Space-Efficient Flash Caching via Compression-before-DeduplicationabstractLarge-scale storage systems boost I/O performance via flash caching, but the underlying storage medium of flash caching incurs significant financial costs and also exhibits low endurance. Previous studies adopt compression-after-deduplication to mitigate writing redundant contents into the flash cache, so as to address the cost and endurance issues. However, deduplication and compression have conflicting preferable cases, and compression-after-deduplication essentially compromises the space-saving benefits of either deduplication or compression. To simultaneously preserve the benefits of both approaches, we explore compression-before-deduplication, which applies compression to eliminate byte-level redundancies across data blocks, followed by deduplication to write only a single copy of duplicate compressed blocks into the flash cache. We present CDCache, a space-efficient flash caching system that realizes compression-before-deduplication. It proposes to dynamically adjust the compression range of data blocks, so as to preserve the effectiveness of deduplication on the compressed blocks. Also, it builds on various design techniques to approximately estimate duplicate data blocks and efficiently manage compressed blocks. Trace-driven experiments show that CDCache improves the read hit ratio and the write reduction ratio of a previous compression-after-deduplication approach by up to 1.3× and 1.6×, respectively, while it only has small memory overhead for index management. Hengying Xiao, Jingwei Li 0001, Yanjing Ren, Ruijin Wang, Xiaosong Zhang 0001 |
INFOCOM | 3 |
| 2023 | FeatureSpy: Detecting Learning-Content Attacks via Feature Inspection in Secure Deduplicated Storage
Jingwei Li 0001, Yanjing Ren, Patrick P. C. Lee, Ting Chen 0002, Xiaosong Zhang 0001 |
INFOCOM | 2 |
| 2022 | Revisiting Frequency Analysis against Encrypted Deduplication via Statistical DistributionabstractEncrypted deduplication addresses both security and storage efficiency in large-scale storage systems: it ensures that each plaintext is encrypted to a ciphertext by a symmetric key derived from the content of the plaintext, so as to allow deduplication on the ciphertexts derived from duplicate plaintexts. However, the deterministic nature of encrypted deduplication leaks the frequencies of plaintexts, thereby allowing adversaries to launch frequency analysis against encrypted deduplication and infer the ciphertext-plaintext pairs in storage. In this paper, we revisit the security vulnerability of encrypted deduplication due to frequency analysis, and show that encrypted deduplication can be even more vulnerable to the sophisticated frequency analysis attack that exploits the underlying storage workload characteristics. We propose the distribution-based attack, which builds on a statistical approach to model the relative frequency distributions of plaintexts and ciphertexts, and improves the inference precision (i.e., have high confidence on the correctness of inferred ciphertext-plaintext pairs) of the previous attack. We evaluate the new attack against real-world storage workloads and provide insights into its actual damage. Jingwei Li 0001, Guoli Wei, Jiacheng Liang, Yanjing Ren, Patrick P. C. Lee, Xiaosong Zhang 0001 |
INFOCOM | 4 |
| 2022 | Enabling Secure and Space-Efficient Metadata Management in Encrypted DeduplicationabstractEncrypted deduplication combines encryption and deduplication in a seamless way to provide confidentiality guarantees for the physical data in deduplicated storage, yet it incurs substantial metadata storage overhead due to the additional storage of keys. We present a new encrypted deduplication storage system called${\sf Metadedup}$, which suppresses metadata storage by also applying deduplication to metadata. Its idea builds on indirection, which adds another level of metadata chunks that record metadata information. We find that metadata chunks are highly redundant in real-world workloads and hence can be effectively deduplicated. We further extend${\sf Metadedup}$to incorporate multiple servers via a distributed key management approach, so as to provide both fault-tolerant storage and security guarantees. We extensively evaluate${\sf Metadedup}$from performance and storage efficiency perspectives. We show that${\sf Metadedup}$achieves high throughput in writing and restoring files, and saves the metadata storage by up to 93.94 percent for real-world backup workloads. Jingwei Li 0001, Suyu Huang, Yanjing Ren, Zuoru Yang, Patrick P. C. Lee, Xiaosong Zhang 0001, Yao Hao |
IEEE Trans. Computers | 3 |
| 2022 | Tunable Encrypted Deduplication with Attack-resilient Key ManagementabstractConventional encrypted deduplication approaches retain the deduplication capability on duplicate chunks after encryption by always deriving the key for encryption/decryption from the chunk content, but such a deterministic nature causes information leakage due to frequency analysis. We present TED , a tunable encrypted deduplication primitive that provides a tunable mechanism for balancing the tradeoff between storage efficiency and data confidentiality. The core idea of TED is that its key derivation is based on not only the chunk content but also the number of duplicate chunk copies, such that duplicate chunks are encrypted by distinct keys in a controlled manner. In particular, TED allows users to configure a storage blowup factor, under which the information leakage quantified by an information-theoretic measure is minimized for any input workload. In addition, we extend TED with a distributed key management architecture and propose two attack-resilient key generation schemes that trade between performance and fault tolerance. We implement an encrypted deduplication prototype TEDStore to realize TED in networked environments. Evaluation on real-world file system snapshots shows that TED effectively balances the tradeoff between storage efficiency and data confidentiality, with small performance overhead. Zuoru Yang, Jingwei Li 0001, Yanjing Ren, Patrick P. C. Lee |
ACM Trans. Storage | 3 |
| 2021 | Accelerating Encrypted Deduplication via SGX
Yanjing Ren, Jingwei Li 0001, Zuoru Yang, Patrick P. C. Lee, Xiaosong Zhang 0001 |
USENIX ATC | 1 |
| 2020 | Balancing storage efficiency and data confidentiality with tunable encrypted deduplicationabstractConventional encrypted deduplication approaches retain the deduplication capability on duplicate chunks after encryption by always deriving the key for encryption/decryption from the chunk content, but such a deterministic nature causes information leakage due to frequency analysis. We present TED, a tunable encrypted deduplication primitive that provides a tunable mechanism for balancing the tradeoff between storage efficiency and data confidentiality. The core idea of TED is that its key derivation is based on not only the chunk content but also the number of duplicate chunk copies, such that duplicate chunks are encrypted by distinct keys in a controlled manner. In particular, TED allows users to configure a storage blowup factor, under which the information leakage quantified by an information-theoretic measure is minimized for any input workload. We implement an encrypted deduplication prototype TEDStore to realize TED in networked environments. Evaluation on real-world file system snapshots shows that TED effectively balances the trade-off between storage efficiency and data confidentiality, with small performance overhead. Jingwei Li 0001, Zuoru Yang, Yanjing Ren, Patrick P. C. Lee, Xiaosong Zhang 0001 |
EuroSys | 3 |
| 2019 | Metadedup: Deduplicating Metadata in Encrypted Deduplication via IndirectionabstractEncrypted deduplication combines encryption and deduplication in a seamless way to provide confidentiality guarantees for the physical data in deduplication storage, yet it incurs substantial metadata storage overhead due to the additional storage of keys. We present a new encrypted deduplication storage system called Metadedup, which suppresses metadata storage by also applying deduplication to metadata. Its idea builds on indirection, which adds another level of metadata chunks that record metadata information. We find that metadata chunks are highly redundant in real-world workloads and hence can be effectively deduplicated. In addition, metadata chunks can be protected under the same encrypted deduplication framework, thereby providing confidentiality guarantees for metadata as well. We evaluate Metadedup through microbenchmarks, prototype experiments, and trace-driven simulation. Metadedup has limited computational overhead in metadata processing, and only adds 6.19% of performance overhead on average when storing files in a networked setting. Also, for real-world backup workloads, Metadedup saves the metadata storage by up to 97.46% at the expense of only up to 1.07% of indexing overhead for metadata chunks. Jingwei Li 0001, Patrick P. C. Lee, Yanjing Ren, Xiaosong Zhang 0001 |
MSST | 3 |