VLDB 2026 Research / reviewers in the wild / expert
Rui Zhang 0112
dblp:60/2536-112
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-2989-7662ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic SchedulingabstractExisting large language model (LLM) serving systems typically batch the compute-bound prefill and I/O-bound decoding phases together.This co-location approach not only leads to significant interference between the two phases but also limits resource allocation and placements.To address these limitations, recent work proposes disaggregating the prefill and decoding phases to enhance performance.However, these works often rely on coarse-grained static scheduling strategies, resulting in imbalanced and insufficient resource utilization.For instance, compute resources for the prefill phase may be overloaded while those for the decoding phase remain idle, resulting in performance bottlenecks.In this paper, we propose WindServe, an efficient phase disaggregated LLM serving system that leverages stream-based, finegrained dynamic scheduling to enhance resource utilization and performance.WindServe features a global scheduler that monitors compute and memory resource usage to dynamically orchestrate cross-phase jobs, effectively reducing queuing delay and KV cache swapping overhead.We also introduce a stall-free rescheduling strategy to saturate the memory resources while minimizing the scheduling overhead from KV cache transfers.Furthermore, we design a stream-based approach to mitigate interference between prefill and decoding jobs.Our evaluation demonstrates that Wind-Serve achieves remarkable stability and SLO attainment under highload scenarios, outperforming state-of-the-art phase-disaggregated LLM serving systems by delivering a 4.28× improvement in TTFT median latency and a 1.5× reduction in TPOT P99 latency. Jingqi Feng, Rui Zhang 0112, Sicheng Liang, Ming Yan 0009, Jie Wu 0003 |
ISCA | 3 |
| 2025 | Offloading distributed key-value stores with off-path SmartNICs
Shangyi Sun, Rui Zhang 0112, Ming Yan 0009, Jie Wu 0003 |
Comput. Networks | 2 |
| 2024 | Labor: Adaptive Lazy Compaction for Learned Index in LSM-Tree
Chunpu Huang, Lulu Chen, Rui Zhang 0112, Ming Yan 0009, Jie Wu 0003 |
COCOON (2) | 4 |
| 2024 | HR-Tree: A Hybrid PMem-DRAM and Write-Optimized R-Tree for Spatial Data Storage
Rui Zhang 0112, Lulu Chen, Shangyi Sun, Ming Yan 0009, Jie Wu 0003 |
COCOON (2) | 1 |
| 2024 | Revisiting Learned Index with Byte-addressable Persistent StorageabstractByte-addressable Persistent Storage (BPS), such as persistent memory and CXL-enabled SSDs, has become an extension of main memory. This opens up new possibilities for indexes that operate and persist data directly on the memory bus. Recent learned indexes exploit data distribution and have shown great potential for some workloads. Despite some work proposed for integrating learned indexes into BPS, they are mainly based on Intel’s first-generation persistent memory. The current design suffers from the following problems: 1) Excessive storage line accesses due to large node in learned indexes; 2) Inefficient concurrency control due to volatile cache; 3) Write amplification due to mismatch access granularity. Rui Zhang 0112, Sicheng Liang, Shangyi Sun, Shaonan Ma, Chengying Huan, Lulu Chen, Zhihui Lu 0002, Yang Xu 0010, Ming Yan 0009, Jie Wu 0003 |
ICPP | 1 |
| 2024 | Mitigating Intra-host Network Congestion with SmartNICabstractWith the rapid development and wide deployment of high-speed network technologies like RDMA and the relatively stagnant evolution of intra-host resources, intra-host network congestion has become a potential issue that may affect the QoS of network applications. Offloading hotspot data to modern Smart-NICs, enabling hotspot data access completion on the SmartNIC, and reducing intra-host network traffic, is a promising solution to this issue. However, due to the limited SmartNIC resources and the complexity of network application requirements, achieving efficient offload is challenging.We present Magician, an architecture to mitigate intra-host network congestion with SmartNIC. Magician adopts a client-driven data access approach to avoid performance degradation caused by limited SmartNIC resources. Magician also introduces a SmartNIC-oriented hotspot data update strategy that dynamically refreshes hotspot data with minimal overhead. Moreover, we design a server-centric data consistency mechanism to ensure data consistency under concurrent access. We implement Magician within the key-value store. Evaluation of the key-value store with and without Magician suggests that, in the presence of intra-host network congestion, Magician significantly mitigates intra-host network congestion, leading to improved performance of network applications. Lulu Chen, Chunpu Huang, Rui Zhang 0112, Yiren Zhou, Ming Yan 0009, Jie Wu 0003 |
IWQoS | 4 |
| 2023 | PFtree: Optimizing Persistent Adaptive Radix Tree for PM Systems on eADR Platform
Rui Zhang 0112, Shangyi Sun, Lulu Chen, Yibo Huang 0005, Ming Yan 0009, Jie Wu 0003 |
DASFAA (1) | 1 |
| 2022 | SKV: A SmartNIC-Offloaded Distributed Key-Value StoreabstractIn data center networks, applications such as dis-tributed key-value stores consume a lot of CPU resources. The performance of the entire system drops significantly under heavy load conditions. In order to improve the performance of key-value stores, many existing studies use RDMA (Remote Direct Memory Access) to reduce the communication overhead. However, RDMA primitives can only offload simple operations to the NIC, such as reading and writing remote memory. With the emergence of new hardware like SmartNICs, we consider whether we can offload more complex operations in distributed key-value stores to SmartNICs to reduce the load on CPU. In this paper we present SKV, a SmartNIC-offloaded distributed key-value store. In order to make full use of the offload ability of the SmartNIC, we make a detailed analysis on the characteristic and architecture of SmartNICs and distributed key-value stores. SKV offloads operations such as data replication to the SmartNIC. We design a new replication mechanism, which enables the server to separate background processing from the interaction with clients in the front. We implement SKV on the Mellanox BlueField SmartNIC. Our evaluations show that SKV improves the overall throughput by 14% and reduces latency by 21 % compared with baseline. Shangyi Sun, Rui Zhang 0112, Ming Yan 0009, Jie Wu 0003 |
CLUSTER | 2 |
| 2022 | An ultra-low latency and compatible PCIe interconnect for rack-scale communicationabstractEmerging network-attached resource disaggregation architecture requires ultra-low latency rack-scale communication. However, current hardware offloading (e.g., RDMA) and user-space (e.g., mTCP) communication schemes still rely on heavily layered protocol stacks which requires the translation between PCIe bus and network protocol, or complex connection/memory resource management within RNICs, inevitably bringing latency overhead. Yibo Huang 0005, Ming Yan 0009, Cunming Liang, Yang Xu 0010, Wenxiong Zou, Yiming Zhang 0018, Rui Zhang 0112, Chunpu Huang, Jie Wu 0003 |
CoNEXT | 9 |
| 2020 | BPS: A reliable and efficient pub/sub communication model with blockchain-enhanced paradigm in multi-tenant edge cloud
Yibo Huang 0005, Rui Zhang 0112, Zhihui Lu 0002, Yiming Zhang 0018, Jie Wu 0003, Lu Zhan, Patrick C. K. Hung |
J. Parallel Distributed Comput. | 2 |