EDBT 2026 Demo / reviewers in the wild / expert
Miao Cai 0001
dblp:28/3792-1
· DBLP profile ↗
24ranked-venue papers
15as first author
19since 2021 · last 2026
0000-0003-1707-4025ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 12 first-author · 13 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable RDMA-accelerated Distributed Locks with Shared Stream AbstractionabstractBlazing fast RDMA technology revolutionizes modern distributed systems and propels them to offload performance-critical data paths onto this network fabric. Designing an RDMA-optimized data path needs to clear a main hurdle—non-scalable distributed locks. Through a performance dissection of existing lock schemes, we find that software-based lock request ordering and polling-based lock ownership transfer scale poorly, leading to high NIC contention and heavy network congestion. To resolve these bottlenecks, this paper proposes StreamLock, a scalable lock primitive that co-designs the distributed lock protocol with fast RDMA networks. The core of StreamLock is a novel shared stream abstraction with two mechanisms: (i) scalable request ordering by repurposing the line-speed packet receiving provided by modern NICs; (ii) peer-to-peer notification to achieve one-round-trip-time lock ownership transfer. We implement StreamLock with off-the-shelf RDMA NICs and compare it with state-of-the-art distributed locks. Comprehensive experimental results showcase that StreamLock outperforms them significantly. Miao Cai 0001, Junru Shen, Xiaojian Liao, Rong Gu 0001, Yanchao Zhao, Bing Chen 0002 |
EuroSys | 1 |
| 2026 | Resource Efficiency and Performance Predictability in A Groupwise, Hardware-Prioritized Cache on NVMe SSDs
Miao Cai 0001, Junru Shen |
IEEE Trans. Computers | 1 |
| 2026 | Achieving Both Performance and Reliability in An Asymmetric File System on Disaggregated Persistent MemoryabstractThe ultra-fast persistent memories (PMs) promise a practical solution toward high-performance distributed file systems. This article examines and reveals a cascade of performance and reliability issues in the current PM provision scheme, which not only underutilizes fast PM devices but also leads to severe consequences, such as throughput degradation, load imbalance, and even service outage. To remedy these, we introduce Ethane+, a rack-scale, distributed file system built on disaggregated persistent memory (DPM). Through resource separation using fast data connection technologies, DPM achieves efficient and cost-effective PM sharing while supporting strong fault isolation. To unleash such hardware potentials, Ethane+ incorporates an asymmetric file system architecture inspired by the imbalanced resource provision feature of DPM. It splits a file system into a control-plane FS and a data-plane FS, and designs these two planes with dual goals of best hardware utilization and hardening file system reliability. Evaluation results demonstrate that Ethane+ reaps the DPM hardware benefits, performs up to 60× better than modern distributed file systems, resists both software and hardware faults, and improves data-intensive application throughputs by up to 15×. Miao Cai 0001, Junru Shen |
ACM Trans. Storage | 1 |
| 2025 | MLog: Achieving Low-Latency, Scalable Shared Log Writes via RDMA Multicast Protocol
Yanan Tao, Miao Cai 0001 |
ICA3PP (1) | 2 |
| 2025 | HeatList: The Case for Retrofitting In-memory Range Index with Hotspot AwarenessabstractSurging memory technologies lead to increasing development of in-memory databases and key-value stores. The range index structure serves as the core design component, which is of importance to their performance and efficiency. Although fast memories significantly improve index structure performance at the hardware level, there still exists poorly-explored optimization space due to non-uniform data accesses in realistic applications. This paper extracts and summarizes four common characteristics of hot data, namely small size, rapid shift, bursty traffic, and spatial locality. Then, we design and implement a novel in-memory range index called HeatList to achieve hotspot awareness. Our core idea is separate a hot layer from the compound index structure to tackle hotspot-related challenges independently, yet without affecting cold data performance. We design the hot layer by proposing a variety of novel techniques to extensively optimize the hot data performance and cope with hotspot inherent characteristics. Evaluation results using both synthetic and production workloads demonstrate that HeatList significantly improves index performance by designing a fast path for hot data access. Junru Shen, Miao Cai 0001, Kangyue Gao |
ICPP | 2 |
| 2025 | Scaling Persistent In-Memory Key-Value Stores Over Modern Tiered, Heterogeneous Memory HierarchiesabstractRecent advances in ultra-fast non-volatile memories (e.g., 3D XPoint) and high-speed interconnect fabrics (e.g., RDMA) enable a high-performance tiered, heterogeneous memory system, effectively overcoming the cost, scaling, and capacity limitations in DRAM-based key-value stores. To fully unleash the performance potential of such memory systems, this paper presents BonsaiKV+, a key-value store that makes the best use of different components in a modern RDMA-enabled heterogeneous memory system. The core of BonsaiKV+ is a tri-layer architecture that achieves efficient, elastic scaling up/out using a set of novel mechanisms and techniques—pipelined tiered indexing, NVM congestion control mechanisms, fine-grained data striping, and NUMA-aware data management—to leverage hardware strengths and tackle device deficiencies. We compare BonsaiKV+ with state-of-the-art key-value stores using a variety of YCSB workloads. Evaluation results demonstrate that BonsaiKV+ outperforms others by up to 7.30$\times$, 18.89$\times$, and 13.67$\times$in read-, write-, and scan-intensive scenarios, respectively. Miao Cai 0001, Junru Shen, Zhihao Qu |
IEEE Trans. Computers | 1 |
| 2024 | Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split LearningabstractThis paper introduces a novel framework designed to achieve a high compression ratio in Split Learning (SL) scenarios where resource-constrained devices are involved in large-scale model training. Our investigations demonstrate that compressing feature maps within SL leads to biased gradients that can negatively impact the convergence rates and diminish the generalization capabilities of the resulting models. Our theoretical analysis provides insights into how compression errors critically hinder SL performance, which previous methodologies underestimate. To address these challenges, we employ a narrow bit-width encoded mask to compensate for the sparsification error without increasing the order of time complexity. Supported by rigorous theoretical analysis, our framework significantly reduces compression errors and accelerates the convergence. Extensive experiments also verify that our method outperforms existing solutions regarding training efficiency and communication complexity. Our code can be found at https://github.com/BinaryMus/MaskSparsification. Zhihao Qu, Shen-Huan Lyu, Miao Cai 0001 |
ECAI | 4 |
| 2024 | RSCache: A Tail Latency Friendly Cache Based on NVMe SSDsabstractHigh fan-out requests are prevalent in systems employing multi-tier architectures. These requests are divided into several sub-requests for parallel processing. However, a high fan-out request must await all sub-requests to be completed before returning, but the processing times of sub-requests are unpredictable due to their differences in characteristics, such as data volume and data popularity. Meanwhile, existing SSD-based caches struggle to adjust request processing speeds to ensure timely handling. As a result, some sub-requests are delayed, affecting the overall latency and causing long tail latency issues.This paper proposes RSCache, a tail latency-friendly cache based on NVMe SSDs. RSCache combines the NVMe Weighted Round-robin (WRR) arbitration mechanism with a priority-based scheduling mechanism to enable differentiated request processing. We propose a fan-out size-based priority assignment strategy along with a latency-aware sub-request sorting method. They collaborate to prioritize sub-requests according to their impacts on tail latency and schedule them to NVMe priority queues with various processing speeds, effectively reducing the processing time variations among sub-requests. In addition, we balance the load across queues to avoid congestion with a novel feedback mechanism. We implement the prototype of RSCache based on SPDK. Our experiments demonstrate that RSCache reduces tail latency by up to 40% and improves throughput by two times compared to state-of-the-art SSD-based cache designs. Jincheng Lu, Miao Cai 0001 |
ISPA | 2 |
| 2024 | CK-index: A Distribution-Aware Learned Index for Composite KeysabstractThe learned index is a high-performance index structure that uses machine learning methods to predict key positions in a large key space efficiently. Existing learned indexes suffer from underfitting of key-to-position mapping, leading to poor lookup performance. This paper finds that a data distribution property in the widely-used composite key schema addresses this issue effectively. Specifically, the composite key consists of an agglomerate of attributes. Keys with the same attribute value have a regular data distribution, which leads to a higher fitness of key-to-position mapping. Applying the property, we introduce CK-index, a distribution-aware learned index for composite keys. CK-index divides the key space according to attribute values and trains each learned model separately for an attribute to achieve high fitness of key-to-position mapping. Furthermore, it achieves low data storage consumption via storing composite key’s attributes instead of the whole keys. We evaluate the CK-index using real-world datasets. Evaluation results demonstrate that CK-index performs much better in lookup performance, bulk loading time and space consumption compared to B+Tree, RMI, PGM-index and ALEX. Zhengyang Wei, Miao Cai 0001 |
ISPA | 3 |
| 2024 | Ethane: An Asymmetric File System for Disaggregated Persistent Memory
Miao Cai 0001, Junru Shen |
USENIX ATC | 1 |
| 2024 | SplitDB: Closing the Performance Gap for LSM-Tree-Based Key-Value StoresabstractLog Structured Merge Tree (LSM tree) serves as the core data storage engine in modern key-value stores. Its adoption is rapidly accelerated with cloud computing and data center development. Acknowledging its widespread use, the LSM tree still faces severe performance issues such as write stall, write amplification, and read inefficiency. This article presents research on improving LSM-tree-based key-value store performance using emerging Non-Volatile Memory (NVM) technology. Our performance diagnosis reveals that the above-mentioned issues result primarily from intensive hot key-value data processing, which is compounded by slow storage devices. To address hotspot bottlenecks, we propose a split log-structured merge tree over hybrid storage by leveraging the intrinsic hot and cold data separation property of the LSM tree. Our approach promotes frequently accessed, small-sized high levels onto fast NVM and offloads the remaining cold, large-sized low levels into slow devices, effectively closing the performance gap for DRAM-disk-based LSM trees. Additionally, we optimize the split LSM tree read and write performance by proposing a variety of novel techniques. We build a hotspot-aware key-value database named SplitDB and perform extensive experiments. Experimental results demonstrate that SplitDB effectively prevents write stalls, achieves a 6-fold write reduction, and improves read throughputs by 3.5 times compared to state-of-the-art key-value databases. Miao Cai 0001, Xuzhen Jiang, Junru Shen |
IEEE Trans. Computers | 1 |
| 2024 | Exploiting Flat Namespace to Improve File System Metadata Performance on Ultra-Fast, Byte-Addressable NVMsabstractThe conventional file system provides a hierarchical namespace by structuring it as a directory tree. Tree-based namespace structure leads to inefficient file path walk and expensive namespace tree traversal, underutilizing ultra-low access latency and superior sequential performance provided by non-volatile memories (NVMs). This article proposes FlatFS+, an NVM file system that features a flat namespace architecture while providing a compatible hierarchical namespace view. FlatFS+ incorporates three novel techniques: the direct file path walk model, range-optimized B r tree, and compressed index key design with scan and write dual optimization, to fully exploit flat namespace to improve file system metadata performance on ultra-fast, byte-addressable NVMs. Evaluation results demonstrate that FlatFS+ achieves significant performance improvements for metadata-intensive benchmarks and real-world applications compared to other file systems. Miao Cai 0001, Junru Shen, Bin Tang 0002, Hao Huang 0011 |
ACM Trans. Storage | 1 |
| 2023 | BonsaiKV: Towards Fast, Scalable, and Persistent Key-Value Stores with Tiered, Heterogeneous Memory SystemabstractEmerging NUMA/CXL-based tiered memory systems with heterogeneous memory devices such as DRAM and NVMM deliver ultrafast speed, large capacity, and data persistence all at once, offering great promise to high-performance in-memory key-value stores. To fully unleash the performance potential of such memory systems, this paper presents BonsaiKV, a key-value store that makes the best use of different components in a tiered memory system. The core of BonsaiKV is a tri-layer hierarchical storage architecture that separates data indexing, persistence, and scalability from each other and realizes each of them within a specialized software-hardware layer. We design BonsaiKV with a set of novel techniques, including collaborative tiered indexing, NVMM congestion control mechanisms, fine-grained data striping, and NUMA-aware data management, to leverage hardware strengths and tackle device deficiencies. We compare BonsaiKV with state-of-the-art NVMM-optimized key-value stores and persistent index structures using a variety of YCSB workloads. Evaluation results demonstrate that BonsaiKV outperforms others by up to 7.69×, 19.59×, and 12.86× in read-, write- and scan-intensive scenarios, respectively. Miao Cai 0001, Junru Shen, Zhihao Qu |
Proc. VLDB Endow. | 1 |
| 2022 | eSROP Attack: Leveraging Signal Handler to Implement Turing-Complete Attack Under CFI Defense
Tianning Zhang, Miao Cai 0001, Diming Zhang, Hao Huang 0011 |
SecureComm | 2 |
| 2022 | SigGuard: Hardening Vulnerable Signal Handling in Commodity Operating SystemsabstractSignal is a useful mechanism provided by many commodity operating systems. However, current signal handling has serious security concerns due to vulnerable design in missing integrity protections for signal handling control flow. Security weaknesses caused by vulnerable design are exploited by adversaries to mount dangerous control-flow attacks. To tackle these issues, this paper investigates root causes of signal-related attacks and proposes SigGuard to harden vulnerable signal handling mechanism. To protect unsafe signal handler execution flow, we design a customized signal handler CFI framework which supports low-cost, reentrant, online CFI analysis and enforcement. To secure signal handler return control flow, we propose an efficient, software-based, intra-process memory isolation method to ensure signal frame data integrity. We evaluate SigGuard with both security and performance experiments. In security experiments, SigGuard successfully thwarts four signal-based attacks, including two proof-of-concept exploits and two realistic attacks conducted in Nginx and Apache server programs, respectively. We also evaluate SigGuard key techniques with a series of microbenchmarks and real-world applications. Experimental results suggest that key defense techniques used in SigGuard introduce reasonable performance costs. Miao Cai 0001, Junru Shen, Tianning Zhang, Hao Huang 0011 |
SRDS | 1 |
| 2022 | FlatFS: Flatten Hierarchical File System Namespace on Non-volatile Memories
Miao Cai 0001, Junru Shen, Bin Tang 0002, Hao Huang 0011 |
USENIX ATC | 1 |
| 2022 | SeBROP: blind ROP attacks without returns
Tianning Zhang, Miao Cai 0001, Diming Zhang, Hao Huang 0011 |
Frontiers Comput. Sci. | 2 |
| 2022 | FastCache: A write-optimized edge storage system via concurrent merging cache for IoT applications
Lin Qian, Zhihao Qu, Miao Cai 0001, Xiaoliang Wang 0001, Weiguo Duan |
J. Syst. Archit. | 3 |
| 2021 | A survey of operating system support for persistent memory
Miao Cai 0001, Hao Huang 0011 |
Frontiers Comput. Sci. | 1 |
| 2020 | HOOP: Efficient Hardware-Assisted Out-of-Place Update for Non-Volatile MemoryabstractByte-addressable non-volatile memory (NVM) is a promising technology that provides near-DRAM performance with scalable memory capacity. However, it requires atomic data durability to ensure memory persistency. Therefore, many techniques, including logging and shadow paging, have been proposed. However, most of them either introduce extra write traffic to NVM or suffer from significant performance overhead on the critical path of program execution, or even both.In this paper, we propose a transparent and efficient hardware-assisted out-of-place update (HOOP) mechanism that supports atomic data durability, without incurring much extra writes and performance overhead. The key idea is to write the updated data to a new place in NVM, while retaining the old data until the updated data becomes durable. To support this, we develop a lightweight indirection layer in the memory controller to enable efficient address translation and adaptive garbage collection for NVM. We evaluate HOOP with a variety of popular data structures and data-intensive applications, including key-value stores and databases. Our evaluation shows that HOOP achieves low critical-path 1atency with small write amplification, which is close to that of a native system without persistence support. Compared with state-of-the-art crash-consistency techniques, it improves application performance by up to $ 1.7\times$, while reducing the write amplification by up to $ 2.1\times$. HOOP also demonstrates scalable data recovery capability on multi-core systems. Miao Cai 0001, Chance C. Coats, Jian Huang 0006 |
ISCA | 1 |
| 2020 | De-randomizing the Code Segment with Timing Function AttackabstractRecently, many effective defensive methods (e.g., ASLR, execute-only-memory) have been proposed to defeat the code reuse attack in the software system. These approaches provide strong system protection through address randomization or memory access restriction. However, this paper identifies a new weak point in these approaches, i.e., missing time protection. We propose a new attack method called timing function attack, which can initiate a code reuse attack even against the state-of-the-art defense techniques. Previous solutions utilize various techniques to hide the spatial information. However, we still can obtain critical security information through the time channel. Specifically, we leverage the function execution time to conduct a side-channel attack. Further, we de-randomize the code segment layout with the timing-channel attack result. Finally, we perform a code-reuse attack with gadgets gathered in previous steps, compromising the whole system. To validate our timing function attack in the real world, we conduct two attacks on two JavaScript engines, i.e., ChakraCore and Chrome v8. Evaluation results show that our attack can successfully bypass the existing defense techniques, such as function-granularity ASLR and XOM, and escalate the privilege. Besides, we also discuss some solutions to prevent and defend our proposed timing function attack. Tianning Zhang, Miao Cai 0001, Diming Zhang, Hao Huang 0011 |
TrustCom | 2 |
| 2020 | A Scalable Virtual memory system based on decentralization for many-cores
Miao Cai 0001, Diming Zhang, Hao Huang 0011 |
J. Syst. Archit. | 1 |
| 2018 | MedusaVM: Decentralizing Virtual Memory System for Multithreaded Applications on Many-core
Miao Cai 0001, Shenming Liu, Weiyong Yang, Hao Huang 0011 |
ICA3PP (1) | 1 |
| 2017 | tScale: A Contention-Aware Multithreaded Framework for Multicore Multiprocessor SystemsabstractOn the multicore and multiprocessor system, multithreaded applications which are kernel-intensive usually suffer from two kinds of performance issues, first one is frequent context switch between kernel/user mode. Another one is lock contention caused by non-scalable synchronization primitives (e.g., ticket spin lock) and may even result in performance degradation under heavy contention level. Unfortunately, current Linux threading model (i.e., NPTL) which adopts exception-based system call mechanism fails to reduce the excessive system call cost. Besides, conventional threading scheduler which is unconscious of lock contention also lacks the ability to limit the number of system-wide contending parallel threads. Both of them impede the application's throughput increment and may lead to the performance breakdown eventually. In this paper we propose a contention-aware threading framework to alleviate these two problems. Our proposed design is composed of two tightly contected components: system call batching via user-level thread library and a contention-aware scheduler based on non-work-conserving scheduling policy. The user-level threading library gathers multiple system call invocations transparently and deliverys these requests to the underlaying kernel working threads. Therefore, tScale improves application performance by reducing massive context switch cost. Then through continuing monitoring system-wide lock contention level and application's total throughput increment, tScale can quickly adjust the number of contending threads in order to sustain the maximum throughput. The prototype system is implemented on Linux 3.18.30 and Glibc 2.23. In microbenchmarks on a 32-core machine, experiment results show that our approach can not only improve the application throughput by up to 20% but also address the lock contention efficiently. Miao Cai 0001, Shenming Liu, Hao Huang 0011 |
ICPADS | 1 |