VLDB 2026 Research / reviewers in the wild / expert
Xiaoquan Zhang
dblp:38/10663
· DBLP profile ↗
16ranked-venue papers
10as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 5 first-author · 9 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPRINT: Line-Rate In-band Network Telemetry Recovery for Application Optimization
Bingzhen Chen, WaiMing Lau, Xiaoquan Zhang, Fung Po Tso 0001, Lin Cui 0001 |
INFOCOM | 3 |
| 2026 | Monic: In-Network Mixture-of-Experts Inference on Programmable Data Planes
Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Kaimin Wei, Weijia Jia 0001, Lin Cui 0001 |
INFOCOM | 1 |
| 2025 | Planner: A Generative Graph Learning Framework for Noisy and Dynamic In-band Network TelemetryabstractIn-band network telemetry (INT) enables real-time network monitoring by embedding telemetry data into packets. The advent of programmable switches further enhances the flexibility of INT by enabling dynamic customization of telemetry collection at the hardware level. However, the practical application of INT is hampered by significant challenges arising from data noise (due to packet loss, delay, and measurement inaccuracies) and network dynamics (such as changing INT paths and feature requirements). These issues severely degrade the performance of machine learning models used for analyzing INT data, hindering the accurate capture of spatio-temporal network characteristics. This paper presents Planner, a novel generative graph learning framework designed to address these limitations. Planner enables the collection of network features at various levels of granularity on programmable switches. Crucially, it constructs dynamic graphs representing evolving INT paths and employs a hybrid Graph Neural Network (GNN) and Recurrent Neural Network (RNN) architecture to effectively learn spatial and temporal dependencies. Furthermore, Planner incorporates variational inference to generate robust latent representations, mitigating the detrimental effects of noise and instability in INT data. We have implemented a testbed prototype of Planner using Intel Tofino ASIC switches. Extensive experiments demonstrate the performance superiority and robustness of Planner over the baseline methods, achieving a 23.2% improvement in F1 score. Xiaoquan Zhang, Waiming Lau, Lin Cui 0001, Fung Po Tso 0001, Zhuoqian Liang, Zhen Zhang 0017, Yuhui Deng 0001 |
ICNP | 1 |
| 2025 | Quark: Implementing Convolutional Neural Networks Entirely on Programmable Data Plane
Mai Zhang, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Zhen Zhang 0017, Yuhui Deng 0001, Zhetao Li |
INFOCOM | 3 |
| 2025 | Enhancing In-Network Computing Deployment via Collaboration Across PlanesabstractThe new paradigm of In-network computing (INC) permits service computation to be executed within network paths, rather than solely on dedicated servers. Although the programmable data plane has showcased notable performance advantages for INC application deployments, its effectiveness is constrained by resource limitations, potentially impeding the expressiveness and scalability of these deployments. Conversely, delegating computational tasks to the control plane, supported by general-purpose servers with abundant resources, offers increased flexibility. Nonetheless, this strategy compromises efficiency to a considerable extent, particularly when the system operates under heavy load. To simultaneously exploit the efficiency of data plane and the flexibility of control plane, we proposeCarlo, a cross-plane collaborative optimization framework to support the network-wide deployment of multiple INC applications across both the control and data plane.Carlofirst analyzes resource requirements of various INC applications across different planes. It then establishes mathematical models for resource allocation in cross-plane and automatically generates solutions using proposed algorithms. We have implemented the prototype ofCarloon Intel Tofino ASIC switches and DPDK. Experimental results demonstrate thatCarlocan effectively trade off between computation time and deployment performance while avoiding performance degradation. Xiaoquan Zhang, Lin Cui 0001, Waiming Lau, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
IEEE Trans. Computers | 1 |
| 2025 | DisPLOY: Target-Constrained Distributed Deployment for Network Measurement Tasks on Data PlaneabstractIn programmable networks, measurement tasks are placed on programmable switches to monitor network traffic at line rate. These tasks typically require substantial resources (e.g., significant SRAM), while programmable switches are constrained by limited resources due to their hardware design (e.g., Tofino ASIC), making distributed deployment essentially. Measurement tasks must monitor specific network locations or traffic flows, introducing significant complexity in deployment optimization. This target-constrained nature makes task optimization on switches (e.g., task merging) become device-dependent and order-dependent, which can lead to deployment failures or performance degradation if ignored. In this paper, we introduceDisPLOY, a novel target-constrained distributed deployment framework specifically designed for network measurement tasks on the data plane.DisPLOYenables operators to specify monitoring targets—network traffic or device/link—across multiple switches. Given the monitoring targets,DisPLOYeffectively minimizes redundant operations and optimizes deployment to achieve both resource efficiency (e.g., minimizing stage consumption) and high-performance monitoring (e.g., high accuracy). We implement and evaluateDisPLOYthrough deployment on both P4 hardware switches (Intel Tofino ASIC) and BMv2. Experimental results show thatDisPLOYsignificantly reduces stage consumption by up to 66% and improves ARE by up to 78.4% in flow size estimation while maintaining end-to-end performance. Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Monte: SFCs Migration Scheme in the Distributed Programmable Data PlaneabstractService function chains (SFCs) are sequences of network functions that provide specific services to meet operators’ needs in today's ISPs and datacenter networks. To improve the performance of SFCs, programmable data planes are used to leverage their low latency and high performance packet processing. However, SFCs need to be adaptable to dynamics such as changes in requirements and attributes. Therefore, the ability to migrate SFCs is essential. Unfortunately, migrating SFCs in distributed programmable data planes is challenging due to the risk of degraded performance and failure to meet SFCs requirements and resource constraints in switches. In this paper, we proposeMonte, which provides an effective SFCs migration scheme in distributed programmable data planes. We build a novel integer programming model to represent the migration process with constraints on resource limitations of switches and SFCs attributes in the distributed data plane. Additionally, an SFCs migration algorithm is designed to optimize the migration cost by deeply analyzing resource allocation in the switch pipeline.Montehas been implemented on both P4 software switches (Bmv2) and hardware switches (Intel Tofino ASIC). Extensive evaluation results show that the migration cost inMonteis 94.03% lower on average than the state-of-the-art deployment scheme, andMontecan effectively save pipeline resources. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Carlo: Cross-Plane Collaboration for Multiple In-network Computing ApplicationsabstractIn-network computing (INC) is a new paradigm that allows applications to be executed within the network, rather than on dedicated servers. Conventionally, INC applications have been exclusively deployed on the data plane (e.g., programmable ASICs), offering impressive performance capabilities. However, the data plane’s efficiency is hindered by limited resources, which can prevent a comprehensive deployment of applications. On the other hand, offloading compute tasks to the control plane, which is underpinned by general-purpose servers with ample resources, provides greater flexibility. However, this approach comes with the tradeoff of significantly reduced efficiency, especially when the system operates under heavy load. To simultaneously exploit the efficiency of data plane and the flexibility of control plane, we propose Carlo, a cross-plane collaborative optimization framework to support the network-wide deployment of multiple INC applications across both the control and data plane. Carlo first analyzes resource requirements of various INC applications across different planes. It then establishes mathematical models for resource allocation in cross-plane and automatically generates solutions using proposed algorithms. We have implemented the prototype of Carlo on Intel Tofino ASIC switches and DPDK. Experimental results demonstrate that Carlo can compute solutions in a short time while avoiding performance degradation caused by the deployment scheme. Xiaoquan Zhang, Lin Cui 0001, Waiming Lau, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
INFOCOM | 1 |
| 2024 | Enabling locality-sensitive machine learning towards low predictive overhead in flow classification
Wenzhi Li, Lin Cui 0001, Xiaoquan Zhang |
Comput. Networks | 3 |
| 2023 | Compiling Service Function Chains via Fine-Grained Composition in the Programmable Data PlaneabstractService function chains (SFCs) are fundamental services in today's datacenters and ISP networks. Explosive volume of network traffic creates high demands for low latency and high performance. The emergence of programmable data planes has offered a new way to overcome the problem. However, limited by pipeline constraints in hardware architecture, implementing multiple network functions on programmable data planes is challenging. Besides, considering various types of network functions, e.g., stateful network functions, a general model is essential for abstracting distinct network functions. In this article, we proposepSFCwhich provides a fine-grained SFCs deployment scheme in programmable data planes. Control flow graph (CFG) is proposed to abstract and analyze various network functions. Then we model pipeline constraints in the hardware architecture using an ILP (Integer Linear Programming), and model the SFCs deployment in the substrate network as a one big switch (OBS) problem. To reduce deployment cost,pSFCfirst composes multiple SFCs to a compound CFG for eliminating redundant logics within SFCs, further decomposes the compound CFG based on the resource limitation per stage, and finally maps the OBS into the substrate network. We have implementedpSFCin both bmv2 software switch and P4 hardware switch (i.e., Intel Tofino ASIC). Evaluation results show thatpSFCreduces switch costs by 45.7% and decreases average latency by 22% without compromising throughput. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Dapper: Deploying Service Function Chains in the Programmable Data Plane Via Deep Reinforcement LearningabstractNetwork functions perform specific packet processing on network traffic. To meet operators' needs, forming service function chains (SFCs) is a fundamental technique used in today's ISPs and datacenter networks. Implementing SFCs in the programmable data plane with high throughput and low latency is a new approach to satisfy demands of ever-growing network traffic. Previous works have proposed different solutions to solve the problem, but they all inevitably have to make trade-offs between running time and performance. For example, an ILP (Integer Linear Programming) can optimize cost but suffers from long running time in large-scale network topologies. Heuristic algorithms depend strongly on manual designs and usually have a performance gap with the optimal solution. In this paper, we proposeDapper, a framework for deploying SFCs in the programmable data plane using DRL (Deep Reinforcement Learning) with graph convolutional network. In order to expand the searching space to prevent the optimal value from being missed,Dapperallows the RL (Reinforcement Learning) agent to simultaneously extract features from both the substrate network and the hardware pipeline, and exploit a graph convolutional network to enhance performance. Moreover, a mask mechanism is also designed to accelerateDapperand improve its scalability.Dapperhas been implemented and extensively evaluated on both P4 hardware switches (equipped with Intel Tofino ASIC) and software switches (i.e., bmv2). Experimental results show thatDappercan automatically generate deployment solutions in a few seconds of running time after training. They also demonstrate thatDapperreduces hardware stage usage and the latency of SFCs by up to 17.8% and 50$\sim$73% respectively on average when compared with heuristics. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | pSFC: Fine-grained Composition of Service Function Chains in the Programmable Data PlaneabstractDynamic service function chains (SFC) are enabled by network function virtualization on general purpose servers. The emergence of programmable data planes (PDP) has offered a new way for the deployment of SFC. However, the implementation of network functions is constrained by resource limitations in PDPs (e.g., compute and memory resource). Moreover, most of existing works do not consider the optimization of state information (e.g., registers), which is essential for stateful network functions. In this paper, we propose pSFC which provides a fine-grained SFC deployment scheme in the PDP to tackle the problem. We first model network functions as control flow graphs (CFG) and the process of deployment as a one big switch (OBS) problem, and then propose an ILP (Integer Linear Programming) model for resource optimization for the OBS problem, which is NP-hard. To solve this problem efficiently, pSFC first composes multiple SFCs for eliminating redundant resources, decomposes the compound CFG based on the resource limitation per stage, and finally maps OBS into the substrate network. We have implemented pSFC in both bmv2 software switch and P4 hardware switch (i.e., Intel Tofino). Evaluation shows that pSFC reduces switch costs 45.7% and average latency 15% while providing the correctness of the process of SFC. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001 |
CCGRID | 1 |
| 2022 | dDrops: Detecting silent packet drops on programmable data plane
Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001 |
Comput. Networks | 3 |
| 2021 | A survey on stateful data plane in software defined networks
Xiaoquan Zhang, Lin Cui 0001, Kaimin Wei, Fung Po Tso 0001, Yangyang Ji, Weijia Jia 0001 |
Comput. Networks | 1 |
| 2021 | pHeavy: Predicting Heavy Flows in the Programmable Data PlaneabstractSince heavy flows account for a significant fraction of network traffic, being able to predict heavy flows has benefited many network management applications for mitigating link congestion, scheduling of network capacity, exposing network attacks and so on. Existing machine learning based predictors are largely implemented on the control plane of Software Defined Networking (SDN) paradigm. As a result, frequent communication between the control and data planes can cause unnecessary overhead and additional delay in decision making. In this paper, we presentpHeavy, a machine learning based scheme for predicting heavy flows directly on the programmable data plane, thus eliminating network overhead and latency to SDN controller. Considering the scarce memory and limited computation capability in the programmable data plane,pHeavyincludes a packet processing pipeline which deploys pre-trained decision tree models for in-network prediction. We have implementedpHeavyin both bmv2 software switch and P4 hardware switch (i.e., Barefoot Tofino). Evaluation results demonstrate thatpHeavyhas achieved 85% and 98% accuracy after receiving the first 5 and 20 packets of a flow respectively, while being able to reduce the size of decision tree by 5.4x on average. More importantly,pHeavycan predict heavy flows at line rate on the P4 hardware switch. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2015 | Turn Waste into Wealth: On Simultaneous Clustering and Cleaning over Dirty DataabstractDirty data commonly exist. Simply discarding a large number of inaccurate points (as noises) could greatly affect clustering results. We argue that dirty data can be repaired and utilized as strong supports in clustering. To this end, we study a novel problem of clustering and repairing over dirty data at the same time. Referring to the minimum change principle in data repairing, the objective is to find a minimum modification of inaccurate points such that the large amount of dirty data can enhance the clustering. We show that the problem can be formulated as an integer linear programming (ILP) problem. Efficient approximation is then devised by a linear programming (LP) relaxation. In particular, we illustrate that an optimal solution of the LP problem can be directly obtained without calling a solver. A quadratic time approximation algorithm is developed based on the aforesaid LP solution. We further advance the algorithm to linear time cost, where a trade-off between effectiveness and efficiency is enabled. Empirical results demonstrate that both the clustering and cleaning accuracies can be improved by our approach of repairing and utilizing the dirty data in clustering. Shaoxu Song, Chunping Li, Xiaoquan Zhang |
KDD | 3 |