VLDB 2026 Research / reviewers in the wild / expert
Shuai Mu 0001
dblp:98/7677-1
· DBLP profile ↗
20ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-8244-2109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 2 first-author · 5 since 2021Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mako: Speculative Distributed Transactions with Geo-Replication
Weihai Shen, Siddhartha Sen 0001, Sebastian Angel, Shuai Mu 0001 |
OSDI | 5 |
| 2025 | Tiga: Accelerating Geo-Distributed Transactions with Synchronized ClocksabstractThis paper presents Tiga, a new design for geo-replicated and scalable transactional databases such as Google Spanner. Tiga aims to commit transactions within 1 wide-area roundtrip time, or 1 WRTT, for a wide range of scenarios, while maintaining high throughput with minimal computational overhead. Tiga consolidates concurrency control and consensus, completing both strictly serializable execution and consistent replication in a single round. It uses synchronized clocks to proactively order transactions by assigning each a future timestamp at submission. In most cases, transactions arrive at servers before their future timestamps and are serialized according to the designated timestamp, requiring 1 WRTT to commit. In rare cases, transactions are delayed and proactive ordering fails, in which case Tiga falls back to a slow path, committing in 1.5–2 WRTTs. Compared to state-of-the-art solutions, Tiga can commit more transactions at 1-WRTT latency, and incurs much less throughput overhead. Evaluation results show that Tiga outperforms all baselines, achieving 1.3–7.2× higher throughput and 1.4–4.6× lower latency. Tiga is open-sourced at https://github.com/New-Consensus-Concurrency-Control/Tiga. Jinkun Geng, Shuai Mu 0001, Anirudh Sivaraman, Balaji Prabhakar |
SOSP | 2 |
| 2025 | AutoMan: Facilitating Verified Distributed Systems Development Through Automatic Code Generation and Manual Optimizations
Ti Zhou, Christa Jenkins, Omar Chowdhury, Shuai Mu 0001 |
SOSP | 5 |
| 2024 | CausalMesh: A Causal Cache for Stateful Serverless ComputingabstractStateful serverless workflows consist of multiple serverless functions that access state on a remote database. Developers sometimes add a cache layer between the serverless runtime and the database to improve I/O latency. However, in a serverless environment, functions in the same workflow may be scheduled to different nodes with different caches, which can cause non-intuitive anomalies. This paper presents CausalMesh, a novel approach to causally consistent caching in serverless computing. CausalMesh is the first cache system that supports coordination-free and abort-free read/write operations and read transactions when clients roam among multiple servers. CausalMesh also supports read-write transactional causal consistency in the presence of client roaming, but at the cost of abort-freedom. Our evaluation shows that CausalMesh has lower latency and higher throughput than existing proposals. Haoran Zhang 0009, Shuai Mu 0001, Sebastian Angel, Vincent Liu 0001 |
Proc. VLDB Endow. | 2 |
| 2023 | Viper: A Fast Snapshot Isolation CheckerabstractSnapshot isolation (SI) is supported by most commercial databases and is widely used by applications. However, checking SI today---given a set of transactions, checking if they obey SI---is either slow or gives up soundness. Jian Zhang 0102, Ye Ji 0003, Shuai Mu 0001, Cheng Tan 0005 |
EuroSys | 3 |
| 2023 | Waverunner: An Elegant Approach to Hardware Acceleration of State Machine Replication
Mohammadreza Alimadadi, Hieu Mai, Shenghsun Cho, Michael Ferdman, Peter A. Milder, Shuai Mu 0001 |
NSDI | 6 |
| 2023 | NCC: Natural Concurrency Control for Strictly Serializable Datastores by Avoiding the Timestamp-Inversion Pitfall
Haonan Lu, Shuai Mu 0001, Siddhartha Sen 0001, Wyatt Lloyd |
OSDI | 2 |
| 2022 | Rolis: a software approach to efficiently replicating multi-core transactionsabstractThis paper presents Rolis, a new speedy and fault-tolerant replicated multi-core transactional database system. Rolis's aim is to mask the high cost of replication by ensuring that cores are always doing useful work and not waiting for each other or for other replicas. Rolis achieves this by not mixing the multi-core concurrency control with multi-machine replication, as is traditionally done by systems that use Paxos to replicate the transaction commit protocol. Instead, Rolis takes an "execute-replicate-replay" approach. Rolis first speculatively executes the transaction on the leader machine, and then replicates the per-thread transaction log to the followers using a novel protocol that leverages independent Paxos instances to avoid coordination, while still allowing followers to safely replay. The execution, replication, and replay are carefully designed to be scalable and have nearly zero coordination overhead across cores. Our evaluation shows that Rolis can achieve 1.03M TPS (transactions per second) on the TPC-C workload, using a 3-replica setup where each server has 32 cores. This throughput result is orders of magnitude higher than traditional software approaches we tested (e.g., 2PL), and is comparable to state-of-the-art, fault-tolerant, in-memory storage systems built using kernel bypass and advanced networking hardware, even though Rolis runs on commodity machines. Weihai Shen, Ansh Khanna, Sebastian Angel, Siddhartha Sen 0001, Shuai Mu 0001 |
EuroSys | 5 |
| 2022 | DepFast: Orchestrating Code of Quorum Systems
Xuhao Luo, Weihai Shen, Shuai Mu 0001, Tianyin Xu |
USENIX ATC | 3 |
| 2021 | Fail-slow fault tolerance needs programming supportabstractThe need for fail-slow fault tolerance in modern distributed systems is highlighted by the increasingly reported fail-slow hardware/software components that lead to poor performance system-wide. We argue that fail-slow fault tolerance not only needs new distributed protocol designs, but also desires programming support for implementing and verifying fail-slow fault-tolerant code. Our observation is that the inability of tolerating fail-slow faults in existing distributed systems is often rooted in the implementations and is difficult to understand and debug. We designed the Dependably Fast Library (DepFast) for implementing fail-slow tolerant distributed systems. DepFast provides expressive interfaces for taking control of possible fail-slow points in the program to prevent unexpected slowness propagation once and for all. We use DepFast to implement a distributed replicated state machine (RSM) and show that it can tolerate various types of fail-slow faults that affect existing RSM implementations. Andrew Yoo, Yuanli Wang, Ritesh Sinha, Shuai Mu 0001, Tianyin Xu |
HotOS | 4 |
| 2021 | Fault-Tolerant Replication with Pull-Based Consensus in MongoDB
Shuai Mu 0001 |
NSDI | 2 |
| 2020 | Cobra: Making Transactional Key-Value Stores Verifiably Serializable
Cheng Tan 0005, Changgeng Zhao, Shuai Mu 0001, Michael Walfish |
OSDI | 3 |
| 2019 | Deferred Runtime Pipelining for contentious multicore software transactionsabstractDRP is a new concurrency control protocol for software transactional memory that achieves high throughput, even for skewed workloads that exhibit high contention. DRP builds on prior works that chop transactions into pieces to expose more concurrency opportunities, but unlike these works, DRP performs no static analyses and supports arbitrary workloads. DRP achieves a high degree of concurrency across most workloads and guarantees deadlock freedom, strict serializability, and opacity. We incorporate DRP into the software transactional objects library STO [18] and find that DRP improves STO's throughput on several STAMP benchmarks by up to 3.6x. Additionally, an in-memory multicore database implemented with our modified variant of STO outperforms databases that use OCC or transaction chopping for concurrency control. Specifically, DRP achieves 6.6x higher throughput than OCC when contention is high. Compared to transaction chopping, our DRP achieves 3.3x higher throughput when contention is medium or low. Furthermore, our implementation achieves comparable performance to OCC and transaction chopping at other contention levels. Shuai Mu 0001, Sebastian Angel, Dennis E. Shasha |
EuroSys | 1 |
| 2019 | On the Parallels between Paxos and Raft, and how to Port OptimizationsabstractIn recent years, Raft has surpassed Paxos to become the more popular consensus protocol in the industry. While many researchers have observed the similarities between the two protocols, no one has shown how Raft and Paxos are formally related to each other. In this paper, we present a formal mapping between Raft and Paxos, and use this knowledge to port a certain class of optimizations from Paxos to Raft. In particular, our porting method can automatically generate an optimized protocol specification with guaranteed correctness. As case studies, we port and evaluate two optimizations, Mencius and Paxos Quorum Lease to Raft. Changgeng Zhao, Shuai Mu 0001, Haibo Chen 0001, Jinyang Li 0001 |
PODC | 3 |
| 2017 | Giza: Erasure Coding Objects across Global Data Centers
Yu Lin Chen, Shuai Mu 0001, Jinyang Li 0001, Cheng Huang 0002, Jin Li 0001, Aaron Ogus, Douglas Phillips |
USENIX ATC | 2 |
| 2016 | The SNOW Theorem and Latency-Optimal Read-Only Transactions
Haonan Lu, Christopher Hodsdon, Khiem Ngo, Shuai Mu 0001, Wyatt Lloyd |
OSDI | 4 |
| 2016 | Consolidating Concurrency Control and Consensus for Commits under Conflicts
Shuai Mu 0001, Lamont Nelson, Wyatt Lloyd, Jinyang Li 0001 |
OSDI | 1 |
| 2016 | Scaling Multicore Databases via Constrained Parallel ExecutionabstractMulticore in-memory databases often rely on traditional con- currency control schemes such as two-phase-locking (2PL) or optimistic concurrency control (OCC). Unfortunately, when the workload exhibits a non-trivial amount of contention, both 2PL and OCC sacrifice much parallel execution op- portunity. In this paper, we describe a new concurrency control scheme, interleaving constrained concurrency con- trol (IC3), which provides serializability while allowing for parallel execution of certain conflicting transactions. IC3 combines the static analysis of the transaction workload with runtime techniques that track and enforce dependencies among concurrent transactions. The use of static analysis simplifies IC3's runtime design, allowing it to scale to many cores. Evaluations on a 64-core machine using the TPC- C benchmark show that IC3 outperforms traditional con- currency control schemes under contention. It achieves the throughput of 434K transactions/sec on the TPC-C bench- mark configured with only one warehouse. It also scales better than several recent concurrent control schemes that also target contended workloads. Shuai Mu 0001, Han Yi, Haibo Chen 0001, Jinyang Li 0001 |
SIGMOD Conference | 2 |
| 2014 | When paxos meets erasure code: reduce network and storage cost in state machine replicationabstractPaxos-based state machine replication is a key technique to build highly reliable and available distributed services, such as lock servers, databases and other data storage systems. Paxos can tolerate any minority number of node crashes in an asynchronous network environment. Traditionally, Paxos is used to perform a full copy replication across all participants. However, full copy is expensive both in term of network and storage cost, especially in wide area with commodity hard drives. Shuai Mu 0001, Kang Chen 0001, Yongwei Wu 0001 |
HPDC | 1 |
| 2014 | Extracting More Concurrency from Distributed Transactions
Shuai Mu 0001, Wyatt Lloyd, Jinyang Li 0001 |
OSDI | 1 |