EDBT 2026 Demo / reviewers in the wild / expert
Qiaoling Wang
dblp:96/7663
· DBLP profile ↗
11ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Performance RoCE-Capable Multicast for Commodity RDMA DatacentersabstractModern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g.,$5.2\times $faster multicast communication and$2.7\times $higher replication throughput for distributed storage. Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005 |
IEEE Trans. Netw. | 8 |
| 2025 | P4KVS: A Role-Replica Separation Offloading Method to Achieve In-Network Consistency for KV Stores Based on P4 SwitchesabstractStrong consistency, particularly linearizability, is essential for distributed DBMSs deployed in correctness-critical domains such as finance and defense. In general, an optimal linearizability DBMS system focus on two key principles: (1) matching single-node (no-consistency cost) Read/Write performance under strong consistency, and (2) practical deployability via general database compatibility. Unfortunately, existing solutions fall short on both fronts. %However, achieving strong consistency often comes with steep performance penalties. For example, etcd-a widely-used Raft-based system-achieves only ~5.9% of the throughput of LevelDB, a single-node store without consistency overhead. To achieve higher performance, software approaches adopt weaker consistency models (e.g., ZAB), rely on narrow network assumptions (e.g., NOPaxos), or expose protocol internals to clients (e.g., CURP), yet still fail to close the performance gap. Recent programmable networking hardware offers promising advances, yet current hardware solutions face practical limitations, including minimal storage and incompatibility with general-purpose databases. We propose P4KVS, the first practical Raft-based in-network consensus offloading solution leveraging programmable switches (P4) for distributed key-value stores. P4KVS offloads only the Leader role to the switch while retaining Followers on servers. Under linearizability, it achieves 74% of single-node LevelDB's throughput for write-heavy workloads, and up to 222.4% for read-heavy workloads by distributing reads across three replicas. This demonstrates that, even under strong consistency, P4KVS can match or exceed the performance of a single-node system. Compared to etcd (which also uses Raft), P4KVS delivers 37.5× higher read throughput and 3520× lower write latency. These results validate our hardware role-replica separation design in eliminating software Raft bottlenecks, while preserving compatibility via standard database interfaces (e.g., LevelDB, etcd) and scaling beyond typical switch memory constraints. Haojuan Li, Zongpu Zhang, Chenzhen Ye, Ruohan Tang, Jian Li 0021, Haibing Guan, Qiaoling Wang, Pengpeng Zhou |
Proc. ACM Manag. Data | 7 |
| 2024 | Training Job Placement in Clusters with Statistical In-Network AggregationabstractIn-Network Aggregation (INA) offloads the gradient aggregation in distributed training (DT) onto programmable switches, where the switch memory could be allocated to jobs in either synchronous or statistical multiplexing mode. Statistical INA has advantages in switch memory utilization, control-plane simplicity, and management safety, but it faces the problem of cross-layer resource efficiency in job placement. This paper presents a job placement system NetPack for clusters with statistical INA, which aims to maximize the utilization of both computation and network resources. NetPack periodically batches and places jobs into the cluster. When placing a job, NetPack runs a steady state estimation algorithm to acquire the available resources in the cluster, heuristically values each server according to its available resources (GPU and bandwidth), and runs a dynamic programming algorithm to efficiently search for servers with the highest value for the job. Our prototype of NetPack and the experiments demonstrate that NetPack outperforms prior job placement methods by 45% in terms of average job completion time on production traces. Bohan Zhao, Wei Xu 0005, Shuo Liu 0002, Yang Tian 0012, Qiaoling Wang, Wenfei Wu |
ASPLOS (1) | 5 |
| 2024 | Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable MulticastabstractModern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage. Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005 |
HPCA | 8 |
| 2023 | In-Network Aggregation with Transport Transparency for Distributed TrainingabstractRecent In-Network Aggregation (INA) solutions offload the all-reduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation. Shuo Liu 0002, Qiaoling Wang, Junyi Zhang 0005, Wenfei Wu, Qinliang Lin, Yao Liu 0006, Marco Canini, Ray C. C. Cheung, Jianfei He |
ASPLOS (3) | 2 |
| 2023 | Efficient Flow Recording with InheritSketch on Programmable SwitchesabstractSeveral studies have been proposed to deploy the flow recording (i.e., flow size counting and sketching algorithms) on programmable switches for high-speed processing, helping network management tasks like scheduling. Although programmable switches provide a remarkable packet processing speed, they are of compact resources and follow a restrictive pipeline programming. To fit these limitations, current algorithms either sacrifice the recording accuracy or harm the switch throughput. In this paper, we propose InheritSketch for further improvement. InheritSketch utilizes a separation counting fashion, which is memory-efficient for compact switches. It accurately records the more valuable heavy hitters in the large key-value counters (i.e., the primary table), while only sketching non-heavy flows in the small sentinel table. With the recording ongoing, InheritSketch intelligently summarizes the historical recording experience as the basis for flow inheritance. That is, flows with the same IDs as the previous heavy hitters are regarded as new heavy hitters, being recorded in the primary table. To correct some incorrect inheritance, we also propose the flow rebellion, which promotes flows of large sizes but wrongly stored in the sentinel table to the primary table. InheritSketch is also helpful in applications like differentiated scheduling. We compare InheritSketch with six previous recording algorithms on three public traffic datasets, and prototype InheritSketch on a commodity P4 switch. The results demonstrate that InheritSketch reduces the recording errors by at most ∼7×, and that InheritSketch only consumes 10% of hardware resources on the switch. Guorui Xie, Qing Li 0006, Guanglin Duan, Yong Jiang 0001, Zhuyun Qi, Qiaoling Wang |
ICDCS | 7 |
| 2021 | Scalable Fully Pipelined Hardware Architecture for In-Network Aggregated AllReduce CommunicationabstractThe Ring-AllReduce framework is currently the most popular solution to deploy industry-level distributed machine learning tasks. However, only about half of the maximum bandwidth can be achieved in the optimal condition. In recent years, several in-network aggregation frameworks have been proposed to overcome the drawback, but limited hardware information have been disclosed. In this paper, we propose a scalable fully-pipelined architecture that handles tasks like forwarding, aggregation and retransmission with no bandwidth loss. The architecture is implemented on a Xilinx Ultrascale FPGA that connects to 8 working servers with 10 Gb/s network adapters, and it is able to scale to more complicated scenarios involving more workers. Compared with Ring-AllReduce, using AllReduce-Switch improves the efficient bandwidth of AllReduce communication with a ratio of$1.75\times $. In image training tasks, the proposed hardware architecture helps to achieve up to$1.67\times $speedup to the training process. For computing-intensive models, the speedup from communication may be partially hidden by computing. In particular, for ResNet-50, AllReduce-Switch improves the training process with MPI and NCCL by$1.30\times $and$1.04\times $respectively. Yao Liu 0006, Junyi Zhang 0005, Shuo Liu 0002, Qiaoling Wang, Wangchen Dai, Ray C. C. Cheung |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | Adaptive Coherent Sampling for Network Delay MeasurementabstractEnd-to-end network delay, as a metric to indicate the QoS (Quality-of-Service), plays an important role in distributed services. Unfortunately, it is infeasible in practice to know all node-pair delay information due to the quadratic growth of overhead by active probing. In this paper, we leverage the stateof-the-art matrix completion technology for better network delay estimation from limited measurements. Although the number of samples required for exact matrix completion is theoretically bounded, it is practically less helpful as the number cannot be specified. This motivates us to propose an adaptive coherent sampling algorithm to select the elements with larger leverage scores to maintain the characteristic of important rows or columns in the delay matrix. The number of samples is adaptively determined by a proposed stopping criterion. Simulation results based on real-world network delay datasets indicate that our proposed algorithm is capable of providing better performance (improves estimation error by 16.9% and convergence stress by 28.9%) at less cost (reduces number of samples by 3.9% and processing time by 78.6%) than traditionally used algorithms. Shuo Liu 0002, Qiaoling Wang |
ICC | 2 |
| 2019 | Poster: online adaptive sampling for network delay measurement via matrix completionabstractEnd-to-end network delay plays an important role in distributed services. Unfortunately, it is infeasible to know all node-pair delay information in practice due to the quadratic growth of overhead by active probing. In this paper, we leverage the state-of-the-art matrix completion technology for better network delay estimation from limited measurements. Specifically, we formulate the matrix completion problem as a nuclear norm minimization problem which can be solved via convex optimization. We propose an online adaptive sampling strategy for network delay measurement. The key idea is to sample the elements with larger leverage scores to maintain characteristic of important rows or columns of the matrix. The number of samples is adaptively determined by a proposed stopping criterion. A preliminary simulation result based on real-world network delay datasets indicates that our proposed sampling algorithm is capable of providing better performance (smaller estimation error and less convergence stress) at less cost (fewer samples and shorter processing time) than other traditionally used algorithms. Shuo Liu 0002, Qiaoling Wang |
Networking | 2 |
| 2018 | Hybrid holiday traffic predictions in cellular networksabstractNetwork traffic forecasting during holidays is important for efficient congestion management and capacity planning. Unfortunately, there is not always enough relevant historical data to forecast traffic at base stations (BSes) individually - this is true in real systems where BSes may be newly built and the data acquisition centers are updated regularly. Hence, it is preferable to forecast holidays in groups of similar traffic patterns, so that data from different BSes can be gathered together to train a single model, and consistent prediction accuracy can be obtained when the context varies. This paper introduces a decomposed model consisting of trend, seasonality, and holiday components of traffic. Then, we focus on developing the holiday sub-models based on residuals that are not covered by trend and seasonality components. A modified k-means algorithm is proposed to cluster residual holiday data. We evaluate our hybrid holiday traffic prediction algorithm on real cellular network data and compare it with the open-source Prophet model developed by Facebook. The evaluation shows that the hybrid prediction method considerably boosts the performance of holiday predictions. Qiaoling Wang, Qinliang Lin |
NOMS | 2 |
| 2017 | An Empirical Study on Mutation Testing of WS-BPEL ProgramsabstractNowadays, applications are increasingly deployed as Web services in the globally distributed cloud computing environment. Multiple services are normally composed to fulfill complex functionalities. Business Process Execution Language for Web Services (WS-BPEL) is an XML-based service composition language that is used to define a complex business process by orchestrating multiple services. Compared with traditional applications, WS-BPEL programs pose many new challenges to the quality assurance, especially testing, of service compositions. A number of techniques have been proposed for testing WS-BPEL programs, but only a few studies have been conducted to systematically evaluate the effectiveness of these techniques. Mutation testing has been widely acknowledged as not only a testing method in its own right but also a popular technique for measuring the fault-detection effectiveness of other testing methods. Several previous studies have proposed a family of mutation operators for generating mutants by seeding various faults into WS-BPEL programs. In this study, we conduct a series of empirical studies to evaluate the applicability and effectiveness of various mutation operators for WS-BPEL programs. The experimental results provide insightful and comprehensive guidance for mutation testing of WS-BPEL programs in practice. In particular, our work is the systematic study in the selection of effective mutation operators specifically for WS-BPEL programs. Chang-Ai Sun, Qiaoling Wang, Huai Liu, Xiangyu Zhang 0001 |
Comput. J. | 3 |