EDBT 2026 Demo / reviewers in the wild / expert
Shengyun Liu
dblp:52/9705
· DBLP profile ↗
24ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ShadowClone: Accelerating Cross-Shard Transactions via Shadow Accounts
Jiahao Qi, Dian Ding, Feilong Lin, Jie Li 0002, Shengyun Liu, Guangtao Xue, Jiannong Cao 0001 |
ICDCS | 6 |
| 2026 | Integrating 2PC with Consensus for Fast ReplicationabstractFast consensus protocols, such as CURP-Q and NOPaxos, have coupled objectives of reducing latency and tolerating faults, but doing so incurs considerable processing overhead and/or necessitates complex changes to existing architectures. Our insight is that when no adverse execution conditions (including operation conflicts, node failures, and network failures) occur, replication can be done safely without consensus at all. Rather than pursuing one single consensus protocol for fast and fault-tolerant replication, we propose to decouple low latency from fault tolerance, and design a hybrid solution that seamlessly integrates client-coordinated two-phase commit (2PC) with consensus, respectively to reduce latency in normal situations and tolerate faults in faulty situations. In the absence of faults, the system enters the Fast mode, where the client directly broadcasts its requests to all replicas with 2PC completing all updates in the first phase, i.e., in one RTT. Otherwise, the system enters the Consensus mode, which resorts to a consensus protocol to tolerate faults. We have applied the hybrid solution to the widely-used Raft consensus protocol to design xRaft, which can adaptively switch between the Fast and Consensus modes, ensuring correctness with minimal switching overhead. Evaluation shows that xRaft significantly outperforms state-of-the-art consensus protocols. Shengyun Liu, Ruofan Xiong, Tianjing Xu, Yongwei Wu 0001, Yiming Zhang 0003 |
SIGCOMM | 3 |
| 2026 | BIND: Enabling Continuous Transaction Processing During Account Migration in Sharded BlockchainsabstractAccount migration in sharded blockchains presents a critical trade-off between optimization effectiveness and system availability. While dynamically reallocating accounts across shards can significantly reduce cross-shard transaction overhead, existing migration mechanisms cause service disruptions that intensify as state data volumes grow. To address this challenge, we propose BIND, a batch-wise account migration protocol that eliminates service interruptions by enabling continuous transaction processing throughout migration. BIND introduces a dual transaction pool architecture that isolates transactions involving migrating accounts while allowing non-migrating accounts to operate uninterrupted. To optimize migration efficiency, we design a reverse greedy heuristic algorithm that partitions accounts into batches based on community cohesion, maximizing intra-batch connectivity to front-load cross-shard communication reduction. We evaluate BIND using real Ethereum transactions, demonstrating superior performance over existing mechanisms. BIND achieves 12% higher overall throughput, reduces migration time to 23.6%-39.3% of the one-shot baseline (across 1-10Gbps bandwidth), and lowers cross-shard transaction rates by 24.1% compared to random batching. These results confirm BIND as a practical solution for large-scale, non-disruptive account migration in production sharded blockchains. Jiahao Qi, Dian Ding, Jie Li 0002, Jiannong Cao 0001, Yi-Chao Chen 0001, Guangtao Xue, Shengyun Liu |
WWW | 7 |
| 2026 | TopoSegNet: Enhancing Geometric Fidelity of Coastline Extraction via a Joint Segmentation and Topological Reasoning FrameworkabstractCoastline extraction from remote sensing imagery is persistently challenged by intra-class heterogeneity (e.g., diverse coastline types) and boundary ambiguity. Existing methods often exhibit suboptimal performance in complex scenes mixing artificial and natural landforms, as they tend to ignore coastline morphological priors and struggle to recover details in low-contrast regions. To address these issues, this paper introduces TopoSeg-Net, a novel collaborative framework centered on a dual-decoder architecture. A segmentation decoder utilizes a Morphology-Aware Attention (MAA) module to adaptively decouple and model diverse coastline morphologies, and a Structure-Detail Synergistic Enhancement (SDSE) module to reconstruct weak boundaries with high fidelity. Meanwhile, a learnable topology decoder frames topology construction as a graph reasoning task, which ensures the geometric and topological integrity of the final vector output.TopoSegNet was evaluated on the public Landsat-8 and a custom Lianyungang Gaofen-1 (GF-1) dataset. The experimental results show that the proposed method reached 98.64%, 66.80%, and 0.795 on the mIoU, BIoU, and APLS metrics, respectively, verifying its validity and superiority. Compared to state-of-the-art methods, the TopoSegNet model demonstrates significantly higher accuracy and topological fidelity. Binge Cui, Shengyun Liu, Jing Zhang 0163, Yan Lu 0014 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2026 | Customizable Information Dispersal-Based Byzantine Broadcast With Communication Optimization Using Bloom FiltersabstractByzantine Fault-tolerant State Machine Replication (BFT SMR) is essential for ensuring the security of high-level services such as blockchain, particularly in scenarios where a subset of nodes may exhibit arbitrary faults. As Byzantine Fault-tolerant broadcast in asynchronous networks is a core component of BFT SMR, this paper studies this broadcast primitive. We propose a novel broadcast protocol in which nodes encode large messages$M$using erasure codes, aggregate the encoded codewords into a vector via a Bloom filter, and employ error correction codes to encode the vector. This approach achieves the known lower bound on the communication complexity of$O(|M|n+\kappa n^{2})$, where$\kappa$represents the output size of a collision-resistant hash function and$n$is the total number of nodes. Performance tests in the Amazon cloud environment demonstrate that when nodes broadcast large messages ($|M|\gg \kappa n^{2}$), the throughput of the new protocol is twice that of existing solutions. Furthermore, the Bloom filter enables the broadcast protocol to possess customizable features for the first time, allowing developers to reduce communication costs through parameter tuning. Shuangquan Tan, Shengyun Liu |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Cheetah: Metadata Aggregation for Fast Object Storage without Distributed OrderingabstractObject stores usually maintain the mapping of objects to data servers' disk volumes (referred to as volume metadata) in a central directory, while storing the object data's in-volume offsets (referred to as offset metadata) together with the data on data servers. Unfortunately, the separation between volume/offset metadata complicates the processing of an object put: to ensure consistency, the multiple writes of the object's volume/offset metadata and object data have to be orchestrated in a particular order, which severely lowers object I/O performance. We propose a write-optimal structure called MetaX that aggregates all metadata of a put, including both volume and offset metadata as well as other meta information such as data checksum and temporary meta-log. Based on MetaX, we design the Cheetah object store, which organizes object storage into rich metadata storage (on meta servers) and raw data storage (on data servers). Cheetah removes the distributed ordering constraint on the multiple metadata/data writes by enforcing local atomicity of writing MetaX, while still ensuring consistency. Evaluation shows that Cheetah significantly outperforms existing object stores. Yiming Zhang 0003, Li Wang 0123, Shengyun Liu, Shun Gai, Xin Yao 0008, Kai Chen 0005, Dongsheng Li 0001, Jiwu Shu |
EuroSys | 3 |
| 2025 | Monosulfide: A Sharded PoW Blockchain System with Secure Adaptive Mining Power Allocation
Guangtao Xue, Shengyun Liu, Jiahao Qi, Dian Ding |
ICA3PP (6) | 3 |
| 2025 | EquiBFT: A Framework for Achieving Fairness in BFT ConsensusabstractByzantine Fault-Tolerant (BFT) consensus protocols are increasingly utilized in blockchain environments. In such protocols, the leader node holds the authority to dictate the transaction order, potentially impacting the fairness of decentralized finance (DeFi) applications. For instance, attackers can exploit this to manipulate transaction order and conduct front-running attacks. The concept of order-fairness, which recently emerged, has become a critical property for preventing a single node from unilaterally determining transaction order. Protocols designed to uphold order-fairness often rely on the sequence in which transactions appear across the network, a factor that can be influenced by the network’s topology. However, this approach has inherent limitations, such as challenges in avoiding Condorcet cycles (Kelkar et al., Crypto 2020).To address these challenges, we propose a novel definition of fairness that requires concealing transaction content before ordering. Additionally, we extend the definitions of liveness and safety of consensus protocols to cover the transaction decryption process, guaranteeing the successful decryption of transactions. Based on the existing BFT protocol and utilizing threshold encryption algorithms, we designed a framework called EquiBFT which can incorporate fairness to BFT protocols. We have proven that the EquiBFT satisfies fairness while ensuring the liveness and safety. We implemented this framework based on HotStuff (Yin et al., PODC 2019) and validated its feasibility in a real-world network environment. Siwei Cai, Lei Fan 0002, Shengyun Liu, Hong-Sheng Zhou |
ICDCS | 3 |
| 2025 | Vegeta: Enabling Parallel Smart Contract Execution in Leaderless Blockchains
Tianjing Xu, Yongqi Zhong, Yiming Zhang 0003, Ruofan Xiong, Jingjing Zhang 0002, Guangtao Xue, Shengyun Liu |
NSDI | 7 |
| 2025 | Chitu: Avoiding Unnecessary Fallback in Byzantine Consensus
Rongji Huang, Xiangzhe Wang, Xiaofeng Yan, Lei Fan 0002, Guangtao Xue, Shengyun Liu |
USENIX ATC | 6 |
| 2024 | Bandle: Asynchronous State Machine Replication Made EfficientabstractState machine replication (SMR) uses consensus as its core component for reaching agreement among a group of processes, in order to provide fault-tolerant services. Most SMR protocols, such as Paxos and Raft, are designed in the partial synchrony model. Partially synchronous protocols rely on timing assumptions to elect a special role (such as the leader), which may become the performance bottleneck under a heavy workload. From an engineering perspective, partially synchronous protocols have to wait for a pre-defined period of time and implement a (complicated) failover mechanism in order to replace the faulty leader. In contrast, asynchronous protocols are immune to such problems. Bo Wang 0116, Shengyun Liu, Xiangzhe Wang, Wenbo Xu 0002, Jingjing Zhang 0002, Ping Zhong 0002, Yiming Zhang 0003 |
EuroSys | 2 |
| 2024 | Efficient Block Storage in the CloudabstractThis paper presents URSAL, an HDD-only block storage system that achieves ultra-efficiency, reliability, scalability and availability at low cost. Compared to existing block stores such as URSA, Ceph, and Sheepdog, URSAL has the following distinctions. First, since parallelism is harmful to the random I/O performance on HDDs, we restrict URSAL storage servers to conservatively perform parallel I/O on HDDs for avoiding I/O contention and reducing tail latency. Second, URSAL designs a proxy-based storage architecture to separate the high-level and low-level I/O logic, where for each virtual machine (VM) there is one URSAL proxy process running at the client VM side to control (at a high level) the procedure of server-side low-level I/O. Third, to alleviate the problem of low random write performance of HDDs, URSAL selectively performs direct block writes on raw HDDs or indirect log appends to HDD journals (which are then asynchronously replayed to raw HDDs), depending on the characteristics of the workloads. Fourth, software failures are nontrivial in large-scale block storage systems of which the availability is vital to client VMs, and thus for high availability we design an efficient fault-tolerance mechanism by isolating the connection management module of URSAL proxy. We have implemented URSAL and deployed it at scale. Extensive evaluation results demonstrate that URSAL achieves much higher performance than the state-of-the-art solutions for underloaded scenarios. Yiming Zhang 0003, Huiba Li, Ping Zhong 0002, Shengyun Liu, Dongsheng Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Bridging the Gap of Timing Assumptions in Byzantine ConsensusabstractAsynchronous Byzantine Fault-Tolerant (BFT) consensus protocols maintain strong consistency across nodes (i.e., ensure safety) and terminate probabilistically (i.e., ensure liveness) despite unbounded network delay. In contrast to protocols under partial synchrony, asynchronous counterparts pay no extra timing assumptions for electing a special role, and thus is more robust to network issues. To formally study this feature, we propose a new classification method for consensus and accordingly categorize relevant work: timing-balanced protocols are those that do not introduce strictly stronger timing-related assumptions for liveness, compared to ones required by safety. Lei Fan 0002, Shengyun Liu, Marko Vukolic, Xiangzhe Wang, Jingjing Zhang 0002 |
Middleware | 3 |
| 2023 | Flexible Advancement in Asynchronous BFT ConsensusabstractByzantine fault tolerant (BFT) consensus protocols are becoming an appealing solution to blockchains. As most blockchain systems are deployed on Wide Area Networks (WANs), with each node acting on behalf of its entity, partially synchronous BFT protocols that rely on network synchrony to elect a single leader can be ill-suited. In contrast, asynchronous protocols have no such timing assumptions. Existing asynchronous protocols confront challenges in terms of both flexibility and performance. Shengyun Liu, Wenbo Xu 0002, Chen Shan, Xiaofeng Yan, Tianjing Xu, Bo Wang 0116, Lei Fan 0002, Fuxi Deng, Ying Yan 0002, Hui Zhang 0002 |
SOSP | 1 |
| 2023 | Hybrid Block Storage for Efficient Cloud Volume ServiceabstractThe migration of traditional desktop and server applications to the cloud brings challenge of high performance, high reliability, and low cost to the underlying cloud storage. To satisfy the requirement, this article proposes a hybrid cloud-scale block storage system called Ursa . Trace analysis shows that the I/O patterns served by block storage have only limited locality to exploit. Therefore, instead of using solid state drives (SSDs) as a cache layer, Ursa proposes hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on hard disk drives (HDDs) . At the core of Ursa ’s hybrid storage design is an adaptive journal that can bridge the performance gap between primary SSDs and backup HDDs for random writes by transforming small backup writes into journal appends, which are then asynchronously replayed and merged to backup HDDs. To efficiently index the journal, we design a novel range-optimized merge-tree structure that combines a continuous range of keys into a single composite key {offset,length} . Ursa integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that Ursa in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (IOPS and throughput per core). Yiming Zhang 0003, Huiba Li, Shengyun Liu, Peng Huang 0005 |
ACM Trans. Storage | 3 |
| 2022 | MPCDDI: A Secure Multiparty Computation-Based Deep Learning Framework for Drug-Drug Interaction Predictions
Shengyun Liu, Shaoliang Peng |
ISBRA | 3 |
| 2021 | IAP: Instant Auditing Protocol for Anonymous PaymentsabstractBlockchain(e.g., Bitcoin) has widespread use in digital currency, which is entirely public to all participants, revealing users' privacy and transaction details. Anonymous blockchains without auditing capability can offer strong privacy guarantees, they could be used by illegal activities. However, anonymous blockchains with auditing capability suffer from two limitations: (i) inefficient auditing capability; (ii) lower degree of decentralization. To address these problems, this paper presents IAP, an instant auditing protocol based on anonymous blockchain with strong anonymity guarantees, which uses audit node cluster to implement decentralized instant auditing. The experimental results show that IAP only needs 60 milliseconds to complete an audit on average with 16 audit nodes, which accounts for one thousand of a complete transaction time. IAP can still complete efficient auditing when there are more than half of the nodes are honest in the audit node cluster. Moreover, its performance is virtually unaffected by increased number of transactions. Ping Zhong 0002, Bo Wang 0116, Anning Wang, Yiming Zhang 0003, Shengyun Liu, Qikai Zhong, Xuping Tu |
ICPADS | 5 |
| 2020 | PBS: An Efficient Erasure-Coded Block Storage System Based on Speculative Partial WritesabstractBlock storage provides virtual disks that can be mounted by virtual machines (VMs). Although erasure coding (EC) has been widely used in many cloud storage systems for its high efficiency and durability, current EC schemes cannot provide high-performance block storage for the cloud. This is because they introduce significant overhead to small write operations (which perform partial write to an entire EC group), whereas cloud-oblivious applications running on VMs are often small-write-intensive. We identify the root cause for the poor performance of partial writes in state-of-the-art EC schemes: for each partial write, they have to perform a time-consuming write-after-read operation that reads the current value of the data and then computes and writes the parity delta, which will be used to “patch” the parity in journal replay. In this article, we present a speculative partial write scheme (called P ARI X) that supports fast small writes in erasure-coded storage systems. We transform the original formula of parity calculation to use the data deltas (between the current/original data values), instead of the parity deltas, to calculate the parities in journal replay. For each partial write, this allows P ARI X to speculatively log only the new value of the data without reading its original value. For a series of n partial writes to the same data, P ARI X performs pure write (instead of write-after-read) for the last n -1 ones while only introducing a small penalty of an extra network round-trip time to the first one. Based on P ARI X, we design and implement P ARI X Block Storage (PBS), an efficient block storage system that provides high-performance virtual disk service for VMs running cloud-oblivious applications. PBS not only supports fast partial writes but also realizes efficient full writes, background journal replay, and fast failure recovery with strong consistency guarantees. Both microbenchmarks and trace-driven evaluation show that PBS provides efficient block storage and outperforms state-of-the-art EC-based systems by orders of magnitude. Yiming Zhang 0003, Huiba Li, Shengyun Liu, Guangtao Xue |
ACM Trans. Storage | 3 |
| 2019 | URSA: Hybrid Block Storage for Cloud-Scale Virtual DisksabstractThis paper presents URSA, a hybrid block store that provides virtual disks for various applications to run efficiently on cloud VMs. Trace analysis shows that the I/O patterns served by block storage have limited locality to exploit. Therefore, instead of using SSDs as a cache layer, URSA proposes an SSD-HDD-hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on HDDs, using journals to bridge the performance gap between SSDs and HDDs. URSA integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that URSA in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (performance per core). We also discuss some practical issues in our deployment. Huiba Li, Yiming Zhang 0003, Dongsheng Li 0001, Shengyun Liu, Peng Huang 0005, Zheng Qin 0002, Kai Chen 0005, Yongqiang Xiong |
EuroSys | 5 |
| 2017 | PARIX: Speculative Partial Writes in Erasure-Coded Systems
Huiba Li, Yiming Zhang 0003, Shengyun Liu, Dongsheng Li 0001, Yuxing Peng 0001 |
USENIX ATC | 4 |
| 2017 | Leader Set Selection for Low-Latency Geo-Replicated State MachineabstractModern planetary scale distributed systems largely rely on a State Machine Replication protocol to keep their service reliable, yet it comes with a specific challenge: latency, bounded by the speed of light. In particular, clients of asingle-leaderprotocol, such as Paxos, must communicate with the leader which must in turn communicate with other replicas: inappropriate selection of a leader may result in unnecessary round-trips across the globe. To cope with this limitation, severalall-leaderandleaderlessalternatives have been proposed recently. Unfortunately, none of them fits all circumstances. In this article we argue that the “right” choice of the number of leaders depends on a given replica configuration and the workload. Then we present${\mathsf {Droopy}}$and${\mathsf {Dripple}}$, two sister approaches built upon state machine replication protocols.${\mathsf {Droopy}}$dynamically reconfigures the set of leaders. Whereas,${\mathsf {Dripple}}$coordinates state partitions wisely, so that each partition can be reconfigured (by${\mathsf {Droopy}}$) separately. Our experimental evaluation on Amazon EC2 shows that,${\mathsf {Droopy}}$and${\mathsf {Dripple}}$reduce latency under imbalanced or localized workloads, compared to their native protocol. When most requests are non-commutative, our approaches do not affect the performance of their native protocol and both outperform a state-of-the-art leaderless protocol. Shengyun Liu, Marko Vukolic |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | XFT: Practical Fault Tolerance beyond Crashes
Shengyun Liu, Paolo Viotti, Christian Cachin, Vivien Quéma, Marko Vukolic |
OSDI | 1 |
| 2010 | Automatic Concurrency Management for distributed applicationsabstractBuilding distributed applications is difficult mostly because of concurrency management. Existing approaches primarily include events and threads. Researchers and developers have been debating for decades to prove which is superior. Although the conclusion is far from obvious, this long debate clearly shows that neither of them is perfect. One of the problems is that they are both complex and error-prone. Both events and threads need the programmers to explicitly manage concurrency, and we believe it is just the source of difficulties. In this paper, we propose a novel approach—automatic concurrency management by the runtime system. It dynamically analyzes the programs to discover potential concurrency opportunities; and it dynamically schedules the communication and the computation tasks, resulting in automatic concurrent execution. This approach is inspired by the instruction scheduling technologies used in modern microprocessors, which dynamically exploits instruction-level parallelism. However, hardware scheduling algorithms do not fit software in many aspects, thus we have to design a new scheme completely from scratch. automatic concurrency management is a runtime technique with no modification to the language, compiler or byte code, so it is good at backward compatibility. It is essentially a dynamic optimization for networking programs. Huiba Li, Shengyun Liu, Yuxing Peng 0001, Dongsheng Li 0001 |
ISCC | 2 |
| 2010 | Superscalar communication: A runtime optimization for distributed applications
Huiba Li, Shengyun Liu, Yuxing Peng 0001, Dongsheng Li 0001, Hangjun Zhou, Xicheng Lu |
Sci. China Inf. Sci. | 2 |