Junru Shen

dblp:330/8540 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-9551-0705ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scalable RDMA-accelerated Distributed Locks with Shared Stream Abstraction
abstract
Blazing fast RDMA technology revolutionizes modern distributed systems and propels them to offload performance-critical data paths onto this network fabric. Designing an RDMA-optimized data path needs to clear a main hurdle—non-scalable distributed locks. Through a performance dissection of existing lock schemes, we find that software-based lock request ordering and polling-based lock ownership transfer scale poorly, leading to high NIC contention and heavy network congestion. To resolve these bottlenecks, this paper proposes StreamLock, a scalable lock primitive that co-designs the distributed lock protocol with fast RDMA networks. The core of StreamLock is a novel shared stream abstraction with two mechanisms: (i) scalable request ordering by repurposing the line-speed packet receiving provided by modern NICs; (ii) peer-to-peer notification to achieve one-round-trip-time lock ownership transfer. We implement StreamLock with off-the-shelf RDMA NICs and compare it with state-of-the-art distributed locks. Comprehensive experimental results showcase that StreamLock outperforms them significantly.
Miao Cai 0001, Junru Shen, Xiaojian Liao, Rong Gu 0001, Yanchao Zhao, Bing Chen 0002
EuroSys2
2026 Resource Efficiency and Performance Predictability in A Groupwise, Hardware-Prioritized Cache on NVMe SSDs
Miao Cai 0001, Junru Shen
IEEE Trans. Computers2
2026 Achieving Both Performance and Reliability in An Asymmetric File System on Disaggregated Persistent Memory
abstract
The ultra-fast persistent memories (PMs) promise a practical solution toward high-performance distributed file systems. This article examines and reveals a cascade of performance and reliability issues in the current PM provision scheme, which not only underutilizes fast PM devices but also leads to severe consequences, such as throughput degradation, load imbalance, and even service outage. To remedy these, we introduce Ethane+, a rack-scale, distributed file system built on disaggregated persistent memory (DPM). Through resource separation using fast data connection technologies, DPM achieves efficient and cost-effective PM sharing while supporting strong fault isolation. To unleash such hardware potentials, Ethane+ incorporates an asymmetric file system architecture inspired by the imbalanced resource provision feature of DPM. It splits a file system into a control-plane FS and a data-plane FS, and designs these two planes with dual goals of best hardware utilization and hardening file system reliability. Evaluation results demonstrate that Ethane+ reaps the DPM hardware benefits, performs up to 60× better than modern distributed file systems, resists both software and hardware faults, and improves data-intensive application throughputs by up to 15×.
Miao Cai 0001, Junru Shen
ACM Trans. Storage2
2025 HeatList: The Case for Retrofitting In-memory Range Index with Hotspot Awareness
abstract
Surging memory technologies lead to increasing development of in-memory databases and key-value stores. The range index structure serves as the core design component, which is of importance to their performance and efficiency. Although fast memories significantly improve index structure performance at the hardware level, there still exists poorly-explored optimization space due to non-uniform data accesses in realistic applications. This paper extracts and summarizes four common characteristics of hot data, namely small size, rapid shift, bursty traffic, and spatial locality. Then, we design and implement a novel in-memory range index called HeatList to achieve hotspot awareness. Our core idea is separate a hot layer from the compound index structure to tackle hotspot-related challenges independently, yet without affecting cold data performance. We design the hot layer by proposing a variety of novel techniques to extensively optimize the hot data performance and cope with hotspot inherent characteristics. Evaluation results using both synthetic and production workloads demonstrate that HeatList significantly improves index performance by designing a fast path for hot data access.
Junru Shen, Miao Cai 0001, Kangyue Gao
ICPP1
2025 Scaling Persistent In-Memory Key-Value Stores Over Modern Tiered, Heterogeneous Memory Hierarchies
abstract
Recent advances in ultra-fast non-volatile memories (e.g., 3D XPoint) and high-speed interconnect fabrics (e.g., RDMA) enable a high-performance tiered, heterogeneous memory system, effectively overcoming the cost, scaling, and capacity limitations in DRAM-based key-value stores. To fully unleash the performance potential of such memory systems, this paper presents BonsaiKV+, a key-value store that makes the best use of different components in a modern RDMA-enabled heterogeneous memory system. The core of BonsaiKV+ is a tri-layer architecture that achieves efficient, elastic scaling up/out using a set of novel mechanisms and techniques—pipelined tiered indexing, NVM congestion control mechanisms, fine-grained data striping, and NUMA-aware data management—to leverage hardware strengths and tackle device deficiencies. We compare BonsaiKV+ with state-of-the-art key-value stores using a variety of YCSB workloads. Evaluation results demonstrate that BonsaiKV+ outperforms others by up to 7.30$\times$, 18.89$\times$, and 13.67$\times$in read-, write-, and scan-intensive scenarios, respectively.
Miao Cai 0001, Junru Shen, Zhihao Qu
IEEE Trans. Computers2
2024 Ethane: An Asymmetric File System for Disaggregated Persistent Memory
Miao Cai 0001, Junru Shen
USENIX ATC2
2024 SplitDB: Closing the Performance Gap for LSM-Tree-Based Key-Value Stores
abstract
Log Structured Merge Tree (LSM tree) serves as the core data storage engine in modern key-value stores. Its adoption is rapidly accelerated with cloud computing and data center development. Acknowledging its widespread use, the LSM tree still faces severe performance issues such as write stall, write amplification, and read inefficiency. This article presents research on improving LSM-tree-based key-value store performance using emerging Non-Volatile Memory (NVM) technology. Our performance diagnosis reveals that the above-mentioned issues result primarily from intensive hot key-value data processing, which is compounded by slow storage devices. To address hotspot bottlenecks, we propose a split log-structured merge tree over hybrid storage by leveraging the intrinsic hot and cold data separation property of the LSM tree. Our approach promotes frequently accessed, small-sized high levels onto fast NVM and offloads the remaining cold, large-sized low levels into slow devices, effectively closing the performance gap for DRAM-disk-based LSM trees. Additionally, we optimize the split LSM tree read and write performance by proposing a variety of novel techniques. We build a hotspot-aware key-value database named SplitDB and perform extensive experiments. Experimental results demonstrate that SplitDB effectively prevents write stalls, achieves a 6-fold write reduction, and improves read throughputs by 3.5 times compared to state-of-the-art key-value databases.
Miao Cai 0001, Xuzhen Jiang, Junru Shen
IEEE Trans. Computers3
2024 Exploiting Flat Namespace to Improve File System Metadata Performance on Ultra-Fast, Byte-Addressable NVMs
abstract
The conventional file system provides a hierarchical namespace by structuring it as a directory tree. Tree-based namespace structure leads to inefficient file path walk and expensive namespace tree traversal, underutilizing ultra-low access latency and superior sequential performance provided by non-volatile memories (NVMs). This article proposes FlatFS+, an NVM file system that features a flat namespace architecture while providing a compatible hierarchical namespace view. FlatFS+ incorporates three novel techniques: the direct file path walk model, range-optimized B r tree, and compressed index key design with scan and write dual optimization, to fully exploit flat namespace to improve file system metadata performance on ultra-fast, byte-addressable NVMs. Evaluation results demonstrate that FlatFS+ achieves significant performance improvements for metadata-intensive benchmarks and real-world applications compared to other file systems.
Miao Cai 0001, Junru Shen, Bin Tang 0002, Hao Huang 0011
ACM Trans. Storage2
2023 BonsaiKV: Towards Fast, Scalable, and Persistent Key-Value Stores with Tiered, Heterogeneous Memory System
abstract
Emerging NUMA/CXL-based tiered memory systems with heterogeneous memory devices such as DRAM and NVMM deliver ultrafast speed, large capacity, and data persistence all at once, offering great promise to high-performance in-memory key-value stores. To fully unleash the performance potential of such memory systems, this paper presents BonsaiKV, a key-value store that makes the best use of different components in a tiered memory system. The core of BonsaiKV is a tri-layer hierarchical storage architecture that separates data indexing, persistence, and scalability from each other and realizes each of them within a specialized software-hardware layer. We design BonsaiKV with a set of novel techniques, including collaborative tiered indexing, NVMM congestion control mechanisms, fine-grained data striping, and NUMA-aware data management, to leverage hardware strengths and tackle device deficiencies. We compare BonsaiKV with state-of-the-art NVMM-optimized key-value stores and persistent index structures using a variety of YCSB workloads. Evaluation results demonstrate that BonsaiKV outperforms others by up to 7.69×, 19.59×, and 12.86× in read-, write- and scan-intensive scenarios, respectively.
Miao Cai 0001, Junru Shen, Zhihao Qu
Proc. VLDB Endow.2
2022 SigGuard: Hardening Vulnerable Signal Handling in Commodity Operating Systems
abstract
Signal is a useful mechanism provided by many commodity operating systems. However, current signal handling has serious security concerns due to vulnerable design in missing integrity protections for signal handling control flow. Security weaknesses caused by vulnerable design are exploited by adversaries to mount dangerous control-flow attacks. To tackle these issues, this paper investigates root causes of signal-related attacks and proposes SigGuard to harden vulnerable signal handling mechanism. To protect unsafe signal handler execution flow, we design a customized signal handler CFI framework which supports low-cost, reentrant, online CFI analysis and enforcement. To secure signal handler return control flow, we propose an efficient, software-based, intra-process memory isolation method to ensure signal frame data integrity. We evaluate SigGuard with both security and performance experiments. In security experiments, SigGuard successfully thwarts four signal-based attacks, including two proof-of-concept exploits and two realistic attacks conducted in Nginx and Apache server programs, respectively. We also evaluate SigGuard key techniques with a series of microbenchmarks and real-world applications. Experimental results suggest that key defense techniques used in SigGuard introduce reasonable performance costs.
Miao Cai 0001, Junru Shen, Tianning Zhang, Hao Huang 0011
SRDS2
2022 FlatFS: Flatten Hierarchical File System Namespace on Non-volatile Memories
Miao Cai 0001, Junru Shen, Bin Tang 0002, Hao Huang 0011
USENIX ATC2