EDBT 2026 Demo / reviewers in the wild / expert
Yi Liu 0115
dblp:97/4626-115
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-8839-9508ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 2 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
Yifan Hua, Shengze Wang 0007, Ruilin Zhou, Yi Liu 0115, Chen Qian 0001, Xiaoxue Zhang 0001 |
NSDI | 5 |
| 2026 | A Distributed Learned Hash TableabstractDistributed Hash Tables (DHTs) are pivotal in numerous high-impact key-value applications built on distributed networked systems, offering a decentralized architecture that avoids single points of failure and improves data availability. Despite their widespread utility, DHTs face substantial challenges in handling range queries, which are crucial for applications such as LLM serving, distributed storage, databases, content delivery networks, and blockchains. To address this limitation, we present LEAD, a novel system incorporating learned models within DHT structures to significantly optimize range query performance. LEAD utilizes recursive machine learning models as the Learned Hash Function to map and retrieve data across a distributed system while preserving the inherent order of data. LEAD includes the designs to minimize range query latency and message cost while maintaining high scalability and resilience to network churn. Our comprehensive evaluations, conducted in both testbed implementation and simulations, demonstrate that LEAD achieves tremendous advantages in system efficiency compared to existing range query methods in large-scale distributed systems, reducing query latency and message cost by 80% to 90%+. Furthermore, LEAD exhibits scalability and robustness against system churn, providing a robust, scalable structure for efficient data retrieval in distributed key-value systems. Shengze Wang 0004, Yi Liu 0115, Xiaoxue Zhang 0001, Liting Hu, Chen Qian 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | Efficient Vector Search on Disaggregated Memory with d-HNSWabstractEfficient vector query processing is essential for powering large-scale AI applications, such as LLMs. However, existing solutions struggle with growing vector datasets that exceed the memory capacity of a single machine, leading to excessive data movement and resource underutilization in monolithic architectures. Yi Liu 0115, Chen Qian 0001 |
HotStorage | 1 |
| 2025 | CloudQC: A Network-aware Framework for Multi-tenant Distributed Quantum ComputingabstractDistributed quantum computing (DQC) that allows a large quantum circuit to be executed simultaneously on multiple quantum processing units (QPUs) becomes a promising approach to increase the scalability of quantum computing. It is natural to envision the near-future DQC platform as a multi-tenant cluster of QPUs, called a Quantum Cloud. However, no existing DQC work has addressed the two key problems of running DQC in a multi-tenant quantum cloud: placing multiple quantum circuits to QPUs and scheduling network resources to complete these jobs. This work is the first attempt to design a circuit placement and resource scheduling framework for a multi-tenant environment. The proposed framework is called CloudQC, which includes two main functional components, circuit placement and network scheduler, with the objectives of optimizing both quantum network cost and quantum computing time. Experimental results with real quantum circuit workloads show that CloudQC significantly reduces the average job completion time compared to existing DQC placement algorithms for both single-circuit and multi-circuit DQC. We envision this work will motivate more future work on network-aware quantum cloud. Ruilin Zhou, Yuhang Gan, Yi Liu 0115, Chen Qian 0001 |
ICDCS | 3 |
| 2025 | Poster: Vortex: Efficient Decentralized Vector Overlay for Similarity Search and DeliveryabstractNearest-neighbor search over embeddings has become a core primitive for AI and LLM-centric workloads. However, prevailing vector databases remain centralized or cluster-bound, introducing single control points, privacy vulnerabilities, and cost/latency bottlenecks. We present Vortex, a decentralized vector overlay that delivers planet-scale approximate nearest neighbor (ANN) search without a centralized control plane. Vortex integrates three key components: (1) Distributed Learned Hashing (DLH), which collaboratively learns piecewise similarity-preserving hash functions to map semantically related vectors to nearby key ranges while balancing load; (2) a Distributed Hash Table (DHT) for scalable, fault-tolerant routing and churn resilience; and (3) a co-designed Distributed HNSW (DHNSW) index for high-recall, low-latency search on each peer. Preliminary results show that Vortex matches the accuracy and latency of leading centralized systems while reducing per-peer index memory requirements by two orders of magnitude and eliminating any central coordinator—enabling fully decentralized, self-organizing ANN overlay for next-generation AI systems. Shengze Wang 0007, Yi Liu 0115, Chen Qian 0001 |
ICNP | 2 |
| 2025 | A Distributed Learned Hash TableabstractDistributed Hash Tables (DHTs) are pivotal in numerous high-impact key-value applications built on distributed networked systems, offering a decentralized architecture that avoids single points of failure and improves data availability. Despite their widespread utility, DHTs face substantial challenges in handling range queries, which are crucial for applications such as LLM serving, distributed storage, databases, content delivery networks, and blockchains. To address this limitation, we present LEAD, a novel system incorporating learned models within DHT structures to significantly optimize range query performance. LEAD utilizes a recursive machine learning model to map and retrieve data across a distributed system while preserving the inherent order of data. LEAD includes the designs to minimize range query latency and message cost while maintaining high scalability and resilience to network churn. Our comprehensive evaluations, conducted in both testbed implementation and simulations, demonstrate that LEAD achieves tremendous advantages in system efficiency compared to existing range query methods in large-scale distributed systems, reducing query latency and message cost by 80% to 90%+. Furthermore, LEAD exhibits remarkable scalability and robustness against system churn, providing a robust, scalable solution for efficient data retrieval in distributed key-value systems. Shengze Wang 0007, Yi Liu 0115, Xiaoxue Zhang 0001, Liting Hu, Chen Qian 0001 |
ICNP | 2 |
| 2025 | Parrot Hashing: Fast and Low-Memory Table Lookups for Network Applications With One CRC-8abstractKey-value lookup functions have been widely applied to network applications, including FIBs, load balancers, and content distributions. Two key performance requirements of a lookup algorithm are high throughput and small memory cost. One limitation of existing fast network lookup algorithms is that they require multiple independent and uniform hash functions, which cost high computation time and might not be available on existing hardware network devices. Recently developed learned model hashing (LMH) proposes to use a linear machine learning model to replace hash functions to avoid hash computation, but they are not optimized for memory cost. We propose a novel network lookup method called Parrot hashing, which uses a learned model to distribute keys into different buckets and applies a simple perfect hashing method to resolve the collisions of the keys in a bucket. Parrot can be implemented with only one CRC-8, which is available on all network devices. We implement Parrot in three prototypes: a software program on end hosts, a software switch, and a FIB running on a hardware programmable switch. The experimental results show that Parrot achieves the highest lookup throughput on all three prototypes, compared to existing methods. Its memory cost is also significantly lower than that of LMH. Yi Liu 0115, Shouqian Shi, Ruilin Zhou, Yuhang Gan, Chen Qian 0001 |
IEEE Trans. Netw. | 1 |
| 2024 | SpotKV: Improving Read Throughput of KVS by I/O-Aware Cache and Adaptive Cuckoo FiltersabstractLSM tree based stores are a popular database design in modern persistent storage systems due to their efficient writes with sorted keys. However, this hierarchical log structure suffers from extensive read amplification because multiple disk accesses are required when it searches for a key. Recent optimizations of LSM trees propose caching hot keys to reduce I/Os mainly based on their access frequencies. However, our empirical studies show that keys are different in I/O costs, which should also be considered in the caching policy: caching key-value pairs with high I/O cost can effectively improve query latency. In addition, false positives incurred by the Bloom filters in LSM trees introduce a large overhead to access SSTables because the queried keys do not exist. In this work, we design and implement SpotKV, which resolves the above two problems in an LSM tree store by proposing two memory-efficient data structures, weighted Count- Min sketch for access and I/O-aware cache admission and dynamic-seed Cuckoo filters for eliminating false positives, to improve data lookup throughput. We implement SpotKV on Google's LevelDB vl.20. From extensive experimental evaluations, SpotKV achieves 1.2-3.0x read throughput while using the same or smaller memory, compared with several state- of-the-art LSM tree stores under the read-heavy workloads of the YCSB benchmarks. Yi Liu 0115, Ruilin Zhou, Yuhang Gan, Chen Qian 0001 |
CLOUD | 1 |
| 2024 | Poster: Distributed Learned Hash TableabstractDistributed Hash Tables (DHTs) are pivotal in numerous high-impact key-value applications built on distributed networked systems, offering a decentralized architecture that avoids single points of failure and improves data availability. Despite their widespread utility, DHTs face substantial challenges in handling range queries, which are crucial for applications such as storage systems, decentralized databases, content distribution networks, and blockchains. To address this limitation, we present LEAD, a novel system incorporating learned models within DHT structures to significantly optimize range query performance. LEAD utilizes a recursive machine learning model to map and retrieve data across a distributed system while preserving the inherent order of data. Preliminary results indicate LEAD achieves tremendous advantages in system efficiency compared to existing range query methods in large-scale distributed systems while maintaining high scalability and resilience to network churn. Shengze Wang 0007, Yi Liu 0115, Xiaoxue Zhang 0001, Liting Hu, Chen Qian 0001 |
ICNP | 2 |
| 2024 | Scalable, Fast, and Low-Memory Table Lookups for Network Applications With One CRC-8abstractKey-value lookup functions have been widely applied to network applications, including FIBs, load balancers, and content distributions. Two key performance requirements of a lookup algorithm are high throughput and small memory cost. One limitation of existing fast network lookup algorithms is that they require multiple independent and uniform hash functions, which cost high computation time and might not be available on existing hardware network devices. Recently developed learned model hashing (LMH) proposes to use a linear machine learning model to replace hash functions to avoid hash computation, but they are not optimized for memory cost. We propose a novel network lookup method called Parrot hashing, which uses a learned model to distribute keys into different buckets and applies a simple perfect hashing method to resolve the collisions of the keys in a bucket. Parrot can be implemented with only one CRC8, which is available on all network devices. We implement Parrot in three prototypes: a software program on end hosts, a software switch, and a FIB running on a hardware programmable switch. The experimental results show that Parrot achieves the highest lookup throughput on all three prototypes, compared to existing methods. Its memory cost is also significantly lower than that of LMH. Yi Liu 0115, Shouqian Shi, Ruilin Zhou, Yuhang Gan, Chen Qian 0001 |
ICNP | 1 |
| 2024 | Towards Practical Overlay Networks for Decentralized Federated LearningabstractDecentralized federated learning (DFL) uses peer-topeer communication to avoid the single point of failure problem in federated learning and has been considered an attractive solution for machine learning tasks on distributed devices. We provide the first solution to a fundamental network problem of DFL: what overlay network should DFL use to achieve fast training of highly accurate models, low communication, and decentralized construction and maintenance? Overlay topologies of DFL have been investigated, but no existing DFL topology includes decentralized protocols for network construction and topology maintenance. Without these protocols, DFL cannot run in practice. This work presents an overlay network, called FedLay, which provides fast training and low communication cost for practical DFL. FedLay is the first solution for constructing near-random regular topologies in a decentralized manner and maintaining the topologies under node joins and failures. Experiments based on prototype implementation and simulations show that FedLay achieves the fastest model convergence and highest accuracy on real datasets compared to existing DFL solutions while incurring small communication costs and being resilient to node joins and failures. Yifan Hua, Jinlong Pang, Xiaoxue Zhang 0001, Yi Liu 0115, Yang Liu 0018, Chen Qian 0001 |
ICNP | 4 |
| 2024 | Outback: Fast and Communication-efficient Index for Key-Value Store on Disaggregated MemoryabstractDisaggregated memory systems achieve resource utilization efficiency and system scalability by distributing computation and memory resources into distinct pools of nodes. RDMA is an attractive solution to support high-throughput communication between different disaggregated resource pools. However, existing RDMA solutions face a dilemma: one-sided RDMA completely bypasses computation at memory nodes, but its communication takes multiple round trips; two-sided RDMA achieves one-round-trip communication but requires non-trivial computation for index lookups at memory nodes, which violates the principle of disaggregated memory. This work presents Outback, a novel indexing solution for key-value stores with a one-round-trip RDMA-based network that does not incur computation-heavy tasks at memory nodes. Outback is the first to utilize dynamic minimal perfect hashing and separates its index into two components: one memory-efficient and compute-heavy component at compute nodes and the other memory-heavy and compute-efficient component at memory nodes. We implement a prototype of Outback and evaluate its performance in a public cloud. The experimental results show that Outback achieves higher throughput than both the state-of-the-art one-sided RDMA and two-sided RDMA-based in-memory KVS by 1.06--5.03×, due to the unique strength of applying a separated perfect hashing index. Yi Liu 0115, Minghao Xie, Shouqian Shi, Yuanchao Xu 0001, Heiner Litz, Chen Qian 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | EdgeCut: Fast and Low-Overhead Access of User-Associated Contents from Edge ServersabstractUser-associated contents play an increasingly important role in modern network applications. With growing deployments of edge servers, the capacity of content storage in edge clusters significantly increases, which provides great potential to satisfy content requests with much shorter latency. However, the large number of contents also causes the difficulty of searching contents on edge servers in different locations because indexing contents costs huge DRAM on each edge server. In this work, we explore the opportunity of efficiently indexing user-associated contents and propose a scalable content-sharing mechanism for edge servers, called EdgeCut, that significantly reduces content access latency by allowing many edge servers to share their cached contents. We design a compact and dynamic data structure called Ludo Locator that returns the IP address of the edge server that stores the requested user-associated content. We have implemented a prototype of EdgeCut in a real network environment running in a public geo-distributed cloud. The experiment results show that EdgeCut reduces content access latency by up to 50% and reduces cloud traffic by up to 50% compared to existing solutions. The memory cost is less than 50MB for 10 million mobile users. The simulations using real network latency data show EdgeCut's advantages over existing solutions on a large scale. Yi Liu 0115, Minmei Wang, Shouqian Shi, Yang Wang 0009, Chen Qian 0001 |
SEC | 1 |
| 2023 | Concurrent Rate-Adaptive Reading With Passive RFIDsabstractRadio frequency identification (RFID)-assisted management systems have been widely applied in warehousing, logistics, retailing, etc. In these scenarios, RFID-aided applications, e.g., object tracking and human behavior sensing, rely on a high-efficiency tag reading to realize accurate analyses and timely responses. However, serious tag collisions in those large-scale RFID systems will inevitably lead to significant decreases in the tag reading rates. To meet the strict timeliness requirements of those practical applications, we aim to treat the individual reading rate for each item tag differently and focus more attention on those user-interactive ones. However, due to unpredictable user behaviors, it is impractical to infer the user-interactive tags in advance. In addition, keeping focusing on them for continuous monitoring despite user movements and multipath-prevalent environments is also challenging. To solve these problems, we propose Spotlight, the first concurrent rate-adaptive reading system in passive RFIDs. Spotlight screens the ID-agnostic user-interactive tags by proposing a multichannel feature for narrow-band RFID systems without any hardware or protocol modification and achieves rate-adaptive reading by implementing real-time MU-MIMO beamforming. Substantial experiments with 1000+ COTS RFID tags exhibit that Spotlight outperforms the commercial reader by$2.7\times $and the SDR-based reader by$6.12\times $. In addition, Spotlight first proposes the online parallel decoding method to realize concurrency among multiple users, which breaks the commercial protocol’s throughput ceiling (37%) and achieves up to 59% throughputs. Ge Wang 0003, Shouqian Shi, Huazhe Wang, Yi Liu 0115, Chen Qian 0001, Cong Zhao 0006, Wei Xi 0003, Han Ding 0002, Zhiping Jiang, Jizhong Zhao |
IEEE Internet Things J. | 4 |