EDBT 2026 Demo / reviewers in the wild / expert
Xingda Wei
dblp:168/9037
· DBLP profile ↗
30ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0003-4983-6047ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-author · 13 since 2021Software engineering, systems software and programming languages · 8 · 6 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KUNSERVE: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM ServingabstractServing LLMs with a cluster of GPUs is common nowadays, where the serving system must meet strict latency SLOs required by applications. However, the stateful nature of LLM serving requires maintaining huge states (i.e., KVCache) in limited GPU memory. Under spikes in real-world workloads, GPU memory can be easily overloaded, leading to orders of magnitude higher response latency due to queuing introduced by waiting for KVCache to be reclaimed. Prior KVCachecentric approaches handle overloading by dropping, migrating, or swapping KVCache. These methods fail to release sufficient memory quickly with requests still queued. Rongxin Cheng 0001, Yuxin Lai, Xingda Wei, Rong Chen 0001, Haibo Chen 0001 |
EuroSys | 3 |
| 2026 | Fast Cloud Storage for AI Jobs via Grouped I/O API with Transparent Read/Write Optimizations
Yingyi Hao, Ting Yao 0001, Xingda Wei, Dingyan Zhang, Tianle Sun, Zhiyong Fu, Huatao Wu, Rong Chen 0001 |
FAST | 3 |
| 2026 | OneSidedMW: Managing Disaggregated Memory Efficiently, Flexibly, and Securely with RNIC Offloading
Jinyu Gu 0001, Xingda Wei, Yubin Xia |
NSDI | 3 |
| 2026 | MetaAttention: A Unified and Performant Attention Framework across Hardware BackendsabstractComputing attention is the backbone of transformer-based models like large language models. However, the increasing diversity of attention algorithms presents significant challenges for unleashing hardware performance. State-of-the-art variants like FlashAttention target a specific attention algorithm or hardware platform, which fail to generalize to other algorithms and platforms. Yu Cheng 0030, Lei Wang 0222, Yuqing Xia, Ziming Miao, Lingxiao Ma, Fan Yang 0024, Jilong Xue, Zhi Yang 0001, Mao Yang 0004, Xingda Wei, Haibo Chen 0001 |
PPoPP | 11 |
| 2026 | Pegasus: A Data Center Network for Bare-Metal AI CloudabstractToday, AI cloud is key to serving diverse users with AI services, where cloud networking forms the basis. In this paper, we share our experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment. The key designs of Pegasus include: 1) Network virtualization: a DPU-RNIC decoupled collaborative hardware architecture to enable a single DPU to virtualize multiple RNICs while reducing the power consumption. We design two-level flow tables on both DPU and RNICs to support underlay-overlay IP address translation and ensure isolation. For DPU-RNIC communication, we introduce a per-RNIC communication state machine to reduce communication overhead. 2) Network transport: customized and transparent transport offloading in the RNIC for low-latency and high-throughput communication performance for various AI workloads. We carefully offload per-packet load balancing and credit-based congestion control in RNICs, optimizing reorder delay and eliminating the impacts of hardware jitter. Pegasus has been deployed in production for over two years, currently covering 8K GPUs and supporting a wide range of tenants' AI applications. Xianneng Zou, Zhaoxun Zhou, Xingda Wei, Zhaohe Chen, Yinben Xia, Lizhou Gao, Jiajun Liang, Chunxu Zhao, Jiewei Yang, Yunpeng Guan, Dongbo Gu, Chao Pei, Zekun He, Yachen Wang |
SIGCOMM | 9 |
| 2026 | Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized Coordination
Hanze Zhang, Rong Chen 0001, Xingda Wei, Haibo Chen 0001 |
Proc. VLDB Endow. | 4 |
| 2025 | CTXNL: A Software-Hardware Co-designed Solution for Efficient CXL-Based Transaction ProcessingabstractTransaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute Express Link (CXL) as a network alternative, which enables low-latency and cache-coherent shared memory accesses. However, directly adopting standard CXL primitives leads to performance degradation due to the high cost of maintaining cross-node cache coherence. To address the CXL challenges, this paper introduces CTXNL, a software-hardware co-designed system that implements a novel hybrid coherence primitive tailored to the loosely coherent nature of transactional data. The core innovation of CTXNL is empowering transaction system developers with the ability to selectively achieve data coherence. Our evaluations on OLTP workloads demonstrate that CTXNL enhances performance, outperforming current network-based systems and achieves up to 2.08x greater throughput than vanilla CXL memory sharing architectures across universal transaction processing policies. Cong Li 0008, Yijin Guan, Dimin Niu, Tianchan Guan, Zhaoyang Du, Xingda Wei, Guangyu Sun 0003 |
ASPLOS (2) | 8 |
| 2025 | ASDSV: Multimodal Generation Made Efficient with Approximate Speculative Diffusion and Speculative VerificationabstractDiffusion in transformer is central to advances in high-quality multimodal generation
but suffer from high inference latency due to their iterative nature.
Inspired by speculative decoding's success in accelerating large language models,
we propose Approximate Speculative Diffusion with Speculative Verification (ASDSV),
a novel method to enhance the efficiency of diffusion models.
Adapting speculative execution to diffusion processes presents unique challenges.
First, the substantial computational cost of verifying numerous speculative steps
for continuous, high-dimensional outputs makes traditional full verification prohibitively expensive.
Second, determining the optimal number of speculative steps $K$
involves a trade-off between potential acceleration and verification success rates.
To address these, ASDSV introduces two key innovations:
1) A speculative verification technique, which leverages the observed temporal correlation between draft and target model outputs,
efficiently validates $K$ speculative steps by only checking the alignment of the initial and final states, significantly reducing verification overhead.
2) A multi-stage speculative strategy that adjusts $K$ according to the denoising phase—employing smaller $K$ during volatile early stages
and larger $K$ during more stable later stages to optimize the balance between speed and quality.
We apply ASDSV to state-of-the-art diffusion transformers,
including Flux.1-dev for image generation and Wan2.1 for video generation.
Extensive experiments demonstrate that ASDSV achieves up to 1.77$\times$-3.01$\times$ speedup
in model inference with a minimal 0.3\%-0.4\% drop in VBench score,
showcasing its effectiveness in accelerating multimodal diffusion models without significant quality degradation.
The code will be publicly available once the acceptance of the paper. Kaijun Zhou 0001, Xingda Wei, Xijun Li, Jinyu Gu 0001 |
NeurIPS | 3 |
| 2025 | ODRP: On-Demand Remote Paging with Programmable RDMA
Xingda Wei, Jinyu Gu 0001, Hongrui Xie, Rong Chen 0001, Haibo Chen 0001 |
NSDI | 2 |
| 2025 | BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching
Dingyan Zhang, Xingda Wei, Yizhou Shan, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 4 |
| 2025 | PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated SpeculationabstractPhoenixOS (PhOS) is the first OS service that can concurrently checkpoint and restore (C/R) GPU processes—a fundamental capability for critical tasks such as fault tolerance, process migration, and fast startup. While concurrent C/R is well-established on CPUs, it poses unique challenges on GPUs due to their lack of essential features for efficiently tracing concurrent memory reads and writes, such as specific hardware capabilities (e.g., dirty bits) and OS-mediated data paths (e.g., copy-on-write). Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen 0001, Mingcong Han, Jinyu Gu 0001, Haibo Chen 0001 |
SOSP | 1 |
| 2025 | KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen 0001, Wenyuan Yu, Haibo Chen 0001 |
USENIX ATC | 3 |
| 2025 | Towards Serialization/Deserialization-free State Transfer in Serverless WorkflowsabstractSerialization and deserialization dominate the state transfer time of serverless workflows, leading to substantial performance penalties when executing various serverless workflow applications. We identify the key reason for serialization and deserialization as a lack of ability to efficiently access the (remote) memory of another function. To this end, we propose RMMap , an OS primitive for remote memory map, which allows a serverless function to directly access the memory of another function, even if it is located remotely. RMMap is the first to completely eliminate serialization and deserialization overhead when transferring states between any pairs of functions in (unmodified) serverless workflows. To make remote memory map efficient and feasible, we co-design it with modern networking (RDMA), OS, language runtime, and serverless platform. Evaluations using real-world serverless workloads show that integrating RMMap with Knative reduces the serverless workflow execution time on Knative by up to 2.6× and improves resource utilizations by 86.3%. Xingda Wei, Fangming Lu, Zhuobin Huang, Rong Chen 0001, Mingyu Wu 0001, Haibo Chen 0001 |
ACM Trans. Comput. Syst. | 1 |
| 2024 | Serialization/Deserialization-free State Transfer in Serverless WorkflowsabstractSerialization and deserialization play a dominant role in the state transfer time of serverless workflows, leading to substantial performance penalties during workflow execution. We identify the key reason as a lack of ability to efficiently access the (remote) memory of another function. We propose RMMap, an OS primitive for remote memory map. It allows a serverless function to directly access the memory of another function, even if it is located remotely. RMMap is the first to completely eliminates serialization and deserialization when transferring states between any pairs of functions in (unmodified) serverless workflows. To make remote memory map efficient and feasible, we co-design it with fast networking (RDMA), OS, language runtime, and serverless platform. Evaluations using real-world serverless workloads show that integrating RMMap with Knative reduces the serverless workflow execution time on Knative by up to 2.6 × and improves resource utilizations by 86.3%. Fangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen 0001, Mingyu Wu 0001, Haibo Chen 0001 |
EuroSys | 2 |
| 2024 | Efficient Distributed File System Offloading on SmartNICabstractModern NIC hardware with programmable capabilities (SmartNIC, SNIC) offers opportunities to offload high-performance distributed file systems (DFS) to the NIC for further acceleration. However, existing designs overlook hardware features and offload all DFS operations to the SNIC, leading to underutilization (18.2%) of SNIC hardware. In this paper, we conduct a systematic study of its inefficiency and identify three non-optimal design and implementation decisions. To address these issues, we propose three improved designs and combine them into S-DFS—an efficient SNIC-offloaded DFS. Compared to the state-of-the-art system, LineFS, S-DFS is 3.18× faster and can fully utilize the power of SNIC. Xingda Wei |
HPCC | 2 |
| 2024 | Locality-Preserving Graph Traversal With Split Live MigrationabstractGraph models many real-world data like social, transportation, biology, and communication data. Hence, graph traversal including multi-hop or graph-walking queries has been the key operation atop graph stores. However, since different graph traversals may touch different sets of vertices, it is hard or even impossible to have a one-size-fits-all graph partitioning algorithm that preserves access locality for various graph traversal workloads. Meanwhile, prior shard-based migration faces a dilemma such that coarse-grained migration may incur more migration overhead over increased locality benefits, while fine-grained migration usually requires excessive metadata and incurs non-trivial maintenance costs. We present Pragh, an efficient locality-preserving live graph migration scheme for graph stores in the form of key-value pairs. The key idea of Pragh is a split migration model that only migrates values physically while retaining keys in the initial location. This allows fine-grained migration while avoiding the need to maintain excessive metadata. Pragh integrates an RDMA-friendly location cache from DrTM-KV to provide fully-localized access to migrated data and further makes a novel reuse of the cache replacement policy for lightweight monitoring. Pragh further supports evolving graphs through a check-and-forward mechanism to resolve the conflict between updates and migration of graph data. Evaluations on an 8-node RDMA-capable cluster (100 Gbps) using a representative graph traversal benchmark show that Pragh can increase the throughput by up to 19× and decrease the median latency by up to 94%, thanks to split live migration that eliminates 97% remote accesses. A port of split live migration to Wukong shows up to 2.53× throughput improvement on representative workloads like LUBM-10240, thanks to a reduction of 88% remote accesses. This further confirms the effectiveness and generality of Pragh. Finally, though Pragh focuses on RDMA-based graph traversal, we show its generality by extending it to support graph traversals under traditional networking. Evaluations on the graph traversal benchmarks and graph query workloads on the same cluster but with 10 Gbps TCP/IP network further confirm its effectiveness without RDMA. Specifically, when evaluating on the LUBM-10240, Wukong-TCP with Pragh can achieve up to 1.87× throughput improvement with a 56% decrease in remote accesses. Rong Chen 0001, Xingda Wei, Xiating Xie, Haibo Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Characterizing Off-path SmartNIC for Accelerating Distributed Systems
Xingda Wei, Rongxin Cheng 0001, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 1 |
| 2023 | No Provisioned Concurrency: Fast RDMA-codesigned Remote Fork for Serverless Computing
Xingda Wei, Fangming Lu, Tianxia Wang, Jinyu Gu 0001, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 1 |
| 2022 | KRCORE: A Microsecond-scale RDMA Control Plane for Elastic Computing
Xingda Wei, Fangming Lu, Rong Chen 0001, Haibo Chen 0001 |
USENIX ATC | 1 |
| 2022 | DrTM+B: Replication-Driven Live Reconfiguration for Fast and General Distributed Transaction ProcessingabstractRecent in-memory database systems leverage advanced hardware features like RDMA to provide transaction processing at millions of transactions per second. Distributed transaction processing systems can scale to even higher rates, especially for partitionable workloads. Unfortunately, it is challenging to sustain such high rates during live reconfiguration of partitions. In this article, we observe that state-of-the-art approaches would cause notable performance disruption under fast transaction processing. To this end, this article presents DrTM+B, a live reconfiguration approach that seamlessly repartitions data with little performance disruption to running transactions. DrTM+B uses a pre-copy-based mechanism to avoid excessive data transfer by leveraging common properties in recent transactional systems. DrTM+B's reconfiguration plans reduce data movement by preferring existing data replicas, while copying data from multiple replicas asynchronously and in parallel. It further reuses the log forwarding mechanism in primary-backup replication to seamlessly track and forward dirty database tuples and avoids iterative copying costs. To commit a reconfiguration plan in a transactional-safe way, DrTM+B designs a cooperative commit protocol for synchronization of data and state among replicas. To boost the performance during data migration, DrTM+B combines the pre-copy and post-copy schemes to propose a hybrid copy scheme. The live reconfiguration approach can also coexist with fault-tolerance mechanisms of primary-backup replication to provide high availability. Evaluation on a working system based on DrTM+R with 3-way replication using typical OLTP workloads like TPC-C and SmallBank shows that DrTM+B incurs only very small performance degradation during live reconfiguration and provides high availability. Both the reconfiguration time and the downtime are also minimal. Sijie Shen, Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Unifying Timestamp with Transaction Ordering for MVCC with Decentralized Scalar Timestamp
Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Zhenhan Gong, Binyu Zang |
NSDI | 1 |
| 2021 | Characterizing and Optimizing Remote Persistent Memory with RDMA and NVM
Xingda Wei, Xiating Xie, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 1 |
| 2021 | XStore: Fast RDMA-Based Ordered Key-Value Store Using Remote Learned CacheabstractRDMA ( Remote Direct Memory Access ) has gained considerable interests in network-attached in-memory key-value stores. However, traversing the remote tree-based index in ordered key-value stores with RDMA becomes a critical obstacle, causing an order-of-magnitude slowdown and limited scalability due to multiple round trips. Using index cache with conventional wisdom—caching partial data and traversing them locally—usually leads to limited effect because of unavoidable capacity misses, massive random accesses, and costly cache invalidations. We argue that the machine learning (ML) model is a perfect cache structure for the tree-based index, termed learned cache . Based on it, we design and implement XStore , an RDMA-based ordered key-value store with a new hybrid architecture that retains a tree-based index at the server to perform dynamic workloads (e.g., inserts) and leverages a learned cache at the client to perform static workloads (e.g., gets and scans). The key idea is to decouple ML model retraining from index updating by maintaining a layer of indirection from logical to actual positions of key-value pairs. It allows a stale learned cache to continue predicting a correct position for a lookup key. XStore ensures correctness using a validation mechanism with a fallback path and further uses speculative execution to minimize the cost of cache misses. Evaluations with YCSB benchmarks and production workloads show that a single XStore server can achieve over 80 million read-only requests per second. This number outperforms state-of-the-art RDMA-based ordered key-value stores (namely, DrTM-Tree, Cell, and eRPC+Masstree) by up to 5.9× (from 3.7×). For workloads with inserts, XStore still provides up to 3.5× (from 2.7×) throughput speedup, achieving 53M reqs/s. The learned cache can also reduce client-side memory usage and further provides an efficient memory-performance tradeoff, e.g., saving 99% memory at the cost of 20% peak throughput. Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
ACM Trans. Storage | 1 |
| 2020 | Fast RDMA-based Ordered Key-Value Store using Remote Learned Cache
Xingda Wei, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 1 |
| 2019 | Pragh: Locality-preserving Graph Traversal with Split Live Migration
Xiating Xie, Xingda Wei, Rong Chen 0001, Haibo Chen 0001 |
USENIX ATC | 2 |
| 2018 | Deconstructing RDMA-enabled Distributed Transactions: Hybrid is Better!
Xingda Wei, Zhiyuan Dong, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 1 |
| 2017 | Replication-driven Live Reconfiguration for Fast Distributed Transaction Processing
Xingda Wei, Sijie Shen, Rong Chen 0001, Haibo Chen 0001 |
USENIX ATC | 1 |
| 2017 | Fast In-Memory Transaction Processing Using RDMA and HTMabstractDrTM is a fast in-memory transaction processing system that exploits advanced hardware features such as remote direct memory access (RDMA) and hardware transactional memory (HTM). To achieve high efficiency, it mostly offloads concurrency control such as tracking read/write accesses and conflict detection into HTM in a local machine and leverages the strong consistency between RDMA and HTM to ensure serializability among concurrent transactions across machines. To mitigate the high probability of HTM aborts for large transactions, we design and implement an optimized transaction chopping algorithm to decompose a set of large transactions into smaller pieces such that HTM is only required to protect each piece. We further build an efficient hash table for DrTM by leveraging HTM and RDMA to simplify the design and notably improve the performance. We describe how DrTM supports common database features like read-only transactions and logging for durability. Evaluation using typical OLTP workloads including TPC-C and SmallBank shows that DrTM has better single-node efficiency and scales well on a six-node cluster; it achieves greater than 1.51, 34 and 5.24, 138 million transactions per second for TPC-C and SmallBank on a single node and the cluster, respectively. Such numbers outperform a state-of-the-art single-node system (i.e., Silo) and a distributed transaction system (i.e., Calvin) by at least 1.9X and 29.6X for TPC-C. Haibo Chen 0001, Rong Chen 0001, Xingda Wei, Jiaxin Shi, Yanzhe Chen, Binyu Zang, Haibing Guan |
ACM Trans. Comput. Syst. | 3 |
| 2016 | Fast and general distributed transactions using RDMA and HTMabstractRecent transaction processing systems attempt to leverage advanced hardware features like RDMA and HTM to significantly boost performance, which, however, pose several limitations like requiring priori knowledge of read/write sets of transactions and providing no availability support. In this paper, we present DrTM+R, a fast in-memory transaction processing system that retains the performance benefit from advanced hardware features, while supporting general transactional workloads and high availability through replication. DrTM+R addresses the generality issue by designing a hybrid OCC and locking scheme, which leverages the strong atomicity of HTM and the strong consistency of RDMA to preserve strict serializability with high performance. To resolve the race condition between the immediate visibility of records updated by HTM transactions and the unready replication of such records, DrTM+R leverages an optimistic replication scheme that uses seqlock-like versioning to distinguish the visibility of tuples and the readiness of record replication. Evaluation using typical OLTP workloads like TPC-C and SmallBank shows that DrTM+R scales well on a 6-node cluster and achieves over 5.69 and 94 million transactions per second without replication for TPC-C and SmallBank respectively. Enabling 3-way replication on DrTM+R only incurs at most 41% overhead before reaching network bottleneck, and is still an order-of-magnitude faster than a state-of-the-art distributed transaction system (Calvin). Yanzhe Chen, Xingda Wei, Jiaxin Shi, Rong Chen 0001, Haibo Chen 0001 |
EuroSys | 2 |
| 2015 | Fast in-memory transaction processing using RDMA and HTMabstractWe present DrTM, a fast in-memory transaction processing system that exploits advanced hardware features (i.e., RDMA and HTM) to improve latency and throughput by over one order of magnitude compared to state-of-the-art distributed transaction systems. The high performance of DrTM are enabled by mostly offloading concurrency control within a local machine into HTM and leveraging the strong consistency between RDMA and HTM to ensure serializability among concurrent transactions across machines. We further build an efficient hash table for DrTM by leveraging HTM and RDMA to simplify the design and notably improve the performance. We describe how DrTM supports common database features like read-only transactions and logging for durability. Evaluation using typical OLTP workloads including TPC-C and SmallBank show that DrTM scales well on a 6-node cluster and achieves over 5.52 and 138 million transactions per second for TPC-C and SmallBank Respectively. This number outperforms a state-of-the-art distributed transaction system (namely Calvin) by at least 17.9X for TPC-C. Xingda Wei, Jiaxin Shi, Yanzhe Chen, Rong Chen 0001, Haibo Chen 0001 |
SOSP | 1 |