Xueying Zhu

dblp:205/7518 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Graph learning · 66% Efficient and distributed learning · 17% Probabilistic and Bayesian machine learning · 17%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 38% Reconfigurable computing and FPGAs · 25% Distributed systems · 25%
Computer networks
3 papers
Software-defined and programmable networks · 91% Internet of things and sensor networks · 9%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software-defined and programmable networks › programmable data plane
programmable switch
1.122025
Hare: A Systematic Framework for Efficient and Generally Automatic Hotspot Offloading on Programmable Switches · IEEE Trans. Netw. 2025
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs · IEEE Trans. Parallel Distributed Syst. 2023
Software-defined and programmable networks
programmable data plane
1.012026
SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC · EuroSys 2026
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload
1.012026
SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC · EuroSys 2026
Machine learning › Graph learning
graph neural network
0.912025
Robust Deep Signed Graph Clustering via Weak Balance Theory · WWW 2025
Machine learning › Graph learning › graph clustering
signed graph clustering
0.912025
Robust Deep Signed Graph Clustering via Weak Balance Theory · WWW 2025
Machine learning › Graph learning › network embedding
signed network embedding
0.912025
Robust Deep Signed Graph Clustering via Weak Balance Theory · WWW 2025
Distributed and cloud data management
query offloading
0.912025
Hare: A Systematic Framework for Efficient and Generally Automatic Hotspot Offloading on Programmable Switches · IEEE Trans. Netw. 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
generalized linear model
0.712023
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs · IEEE Trans. Parallel Distributed Syst. 2023
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.712023
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs · IEEE Trans. Parallel Distributed Syst. 2023
Distributed systems › distributed machine learning
distributed training
0.712023
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs · IEEE Trans. Parallel Distributed Syst. 2023
Reconfigurable computing and FPGAs
FPGA accelerator
0.712023
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs · IEEE Trans. Parallel Distributed Syst. 2023
Internet of things and sensor networks › wireless sensor network
in-network aggregation
0.212023
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs · IEEE Trans. Parallel Distributed Syst. 2023

Methods — techniques the papers use, named apart from their topics

in-cache processing · 2.0header-only offloading · 2.0pipeline parallelism · 2.0model parallelism · 2.0allreduce · 2.0switch-server co-offloading · 1.7MAT-based cross-stage structure · 1.7spectral clustering · 0.9graph neural network · 0.9contrastive learning · 0.9
YearPublicationVenuePosition
2026 SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC
abstract
As the gap between network and CPU speeds rapidly increases, the CPU-centric network stack proves inadequate due to excessive CPU and memory overheads. Though hardware-offloaded network stacks alleviate these issues, they suffer from limited flexibility in both control and data planes. It seems promising to offload network stacks to Smart-NICs to provide high flexibility. However, naive offloading leads to low throughput due to the inherent architectural limitations of widespread off-path SmartNICs. Even simple operations on staged network traffic would overwhelm the limited SmartNIC memory bandwidth. To this end, we design SmartNS, a SmartNIC-centric network stack with software transport programmability and line-rate packet processing capabilities. To tackle the limitations of SmartNIC-induced challenges, we propose a header-only offloading TX path and an unlimited-working-set in-cache processing RX path to minimize memory traffic to fit the wimpy SmartNIC memory bandwidth. To fully utilize the SmartNIC computing resources, we propose a programmable offloading engine to enable cloud providers to offload customized tasks along with the network stack processing. We prototype SmartNS using the widespread Nvidia BlueField-3 SmartNIC, and implement RoCEv2 and Solar transport protocols by leveraging SmartNS's software programmability. SmartNS achieves 2.2× higher throughput than the microkernel-based baseline in block storage disaggregation and 1.3× higher throughput than the hardware-offloaded baseline in KVCache transfer.
Xuzheng Chen, Jie Zhang 0081, Baolin Zhu, Xueying Zhu, Zhongqing Chen, Lingjun Zhu, Yin Zhang 0006, Yuanchao Shu, Peng Cheng 0001, Zeke Wang
EuroSys4
2025 CoTraX: An Efficient Parallel Training Method for On-Policy Deep Reinforcement Learning
Xueying Zhu, Chong Tang 0006
PRICAI (4)3
2025 Robust Deep Signed Graph Clustering via Weak Balance Theory
abstract
Signed graph clustering is a critical technique for discovering community structures in graphs that exhibit both positive and negative relationships. We have identified two significant challenges in this domain: i) existing signed spectral methods are highly vulnerable to noise, which is prevalent in real-world scenarios; ii) the guiding principle "an enemy of my enemy is my friend", rooted in Social Balance Theory, often narrows or disrupts cluster boundaries in mainstream signed graph neural networks. Addressing these challenges, we propose the Deep Signed Graph Clustering framework (DSGC), which leverages Weak Balance Theory to enhance preprocessing and encoding for robust representation learning. First, DSGC introduces Violation Sign-Refine to denoise the signed network by correcting noisy edges with high-order neighbor information. Subsequently, Density-based Augmentation enhances semantic structures by adding positive edges within clusters and negative edges across clusters, following Weak Balance principles. The framework then utilizes Weak Balance principles to develop clustering-oriented signed neural networks to broaden cluster boundaries by emphasizing distinctions between negatively linked nodes. Finally, DSGC optimizes clustering assignments by minimizing a regularized clustering loss. Comprehensive experiments on synthetic and real-world datasets demonstrate DSGC consistently outperforms all baselines, establishing a new benchmark in signed graph clustering.
Xin Li 0033, Zeyu Zhang 0004, Mingzhong Wang, Xueying Zhu, Lejian Liao
WWW5
2025 Hare: A Systematic Framework for Efficient and Generally Automatic Hotspot Offloading on Programmable Switches
abstract
Switch-based hotspot offloading is a trendy solution for latency-sensitive applications to achieve high system throughput with an acceptable P99 query response latency. However, due to the varying object sizes, dynamic workloads, and complex query-processing functions of the latency-sensitive applications, existing switch-based dynamic hotspot offloading approaches struggle to handle these applications effectively. This is mainly because of their inefficient switch resource utilization and non-generalizable hotspot offloading designs. So we propose Hare, a systematic framework that consists of three techniques to address these issues. First, Hare uses a MAT-based cross-stage structure to store and perform hit-checks for large hotspots on the switch data plane. Second, Hare uses a switch-server co-offloading mechanism to support fast and precise offloading. Third, Hare is designed to enable generally automatic offloading by decoupling application-related query processing with hotspot offloading. Compared to the state-of-the-art approaches, Hare supports$8.86\times \sim 9.97\times $larger hotspot size, achieves$1.27 \times \sim 6.61 \times $higher system throughput, and can recover the system throughput and the P99 query response latency within 8s.
Xueying Zhu, Yingtao Li 0001, Xiang Li 0205, Jialin Li 0001, Zeke Wang
IEEE Trans. Netw.1
2024 Staleness-Reduction Mini-Batch K-Means
abstract
K -means (km) is a clustering algorithm that has been widely adopted due to its simple implementation and high clustering quality. However, the standard km suffers from high computational complexity and is therefore time-consuming. Accordingly, the mini-batch (mbatch) km is proposed to significantly reduce computational costs in a manner that updates centroids after performing distance computations on just a mbatch, rather than a full batch, of samples. Even though the mbatch km converges faster, it leads to a decrease in convergence quality because it introduces staleness during iterations. To this end, in this article, we propose the staleness-reduction mbatch (srmbatch) km, which achieves the best of two worlds: low computational costs like the mbatch km and high clustering quality like the standard km. Moreover, srmbatch still exposes massive parallelism to be efficiently implemented on multicore CPUs and many-core GPUs. The experimental results show that srmbatch can converge up to 40× - 130× faster than mbatch when reaching the same target loss, and srmbatch is able to reach 0.2%-1.7% lower final loss than that of mbatch.
Xueying Zhu, Jie Sun 0017, Zhenhao He, Jiantong Jiang, Zeke Wang
IEEE Trans. Neural Networks Learn. Syst.1
2023 P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs
abstract
Generalized linear models (GLMs) are a widely utilized family of machine learning models in real-world applications. As data size increases, it is essential to perform efficient distributed training for these models. However, existing systems for distributed training have a high cost for communication and often use large batch sizes to balance computation and communication, which negatively affects convergence. Therefore, we argue for an efficient distributed GLM training system that strives to achieve linear scalability, while keeping batch size reasonably low. As a start, we propose P4SGD, a distributed heterogeneous training system that efficiently trains GLMs through model parallelism between distributed FPGAs and through forward-communication-backward pipeline parallelism within an FPGA. Moreover, we propose a light-weight, latency-centric in-switch aggregation protocol to minimize the latency of the AllReduce operation between distributed FPGAs, powered by a programmable switch. As such, to our knowledge, P4SGD is the first solution that achieves almost linear scalability between distributed accelerators through model parallelism. We implement P4SGD on eight Xilinx U280 FPGAs and a Tofino P4 switch. Our experiments show P4SGD converges up to 6.5X faster than the state-of-the-art GPU counterpart.
Hongjing Huang, Yingtao Li 0001, Jie Sun 0017, Xueying Zhu, Jie Zhang 0081, Jialin Li 0001, Zeke Wang
IEEE Trans. Parallel Distributed Syst.4
2019 A real-time traffic index model for expressways
abstract
Summary In this paper, a real‐time traffic index model of expressways is proposed by using a traffic index to evaluate the actual conditions of expressways. The model considers the actual situation of floating and nonfloating vehicles on expressways, in the context of massive floating car data. Included is the realization of the complete calculation model of real‐time traffic index estimation, including highway section division, spatial topology map matching, driving route calculation, and road congestion status judgment. For roads without floating car coverage, the weighted‐moving‐average time‐series prediction method is used to predict the traffic index, so that the running condition of all roads in the network can be analyzed completely. The simulation results show that the proposed highway traffic index can reflect not only the overall high‐speed operation but also real‐time congestion at specific high‐speed or specific highway intervals, providing an effective reference for travel.
Fusheng Xu, Zhongxiang Huang, Xueying Zhu, Hongwei Wu, Jun Zhang 0014
Concurr. Comput. Pract. Exp.4