Xuzheng Chen

dblp:367/9362 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0009-7432-5485ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Cloud and datacenter computing · 33% GPUs and heterogeneous computing · 15% Hardware accelerators and domain-specific architectures · 14%
Computer networks
2 papers
Datacenter networks · 53% Software-defined and programmable networks · 47%
Network and information security
1 paper
Authentication and access control · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › network accelerator
SmartNIC
1.622025
RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025
Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024
Software-defined and programmable networks
programmable data plane
1.012026
SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC · EuroSys 2026
Authentication and access control
access control
1.012026
Janus: Enabling Expressive and Efficient ACLs in High-speed RDMA Clouds · NDSS 2026
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload
1.012026
SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC · EuroSys 2026
Datacenter networks › RDMA
RDMA congestion control
0.912025
SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine · USENIX ATC 2025
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems
0.912025
CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025
Cloud and datacenter computing
datacenter network
0.912025
RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025
GPUs and heterogeneous computing
GPU storage access
0.912025
CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025
Distributed systems › remote procedure call
RPC offloading
0.912025
RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025
Performance modeling and evaluation › workload characterization
architectural characterization
0.812024
Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024
Cloud and datacenter computing
datacenter RPC
0.812024
DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024
Distributed systems
remote procedure call
0.812024
DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024
Performance modeling and evaluation
workload characterization
0.812024
Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024
Datacenter networks
RDMA
0.312025
SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine · USENIX ATC 2025
Storage systems › flash and SSD › solid-state drive
NVMe SSD
0.312025
CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025
Interconnection networks and networks-on-chip › high-speed interconnect
PCIe interconnect
0.312025
RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025
Memory systems › memory disaggregation
CXL memory disaggregation
0.212024
DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024
Memory systems
memory disaggregation
0.212024
DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024
Cloud and datacenter computing › computation offloading
network function offloading
0.212024
Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024

Methods — techniques the papers use, named apart from their topics

in-cache processing · 2.0header-only offloading · 2.0RDMA · 2.0target-aware deserializer · 0.9software-programmable congestion control · 0.9memory-affinity CPU-SmartNIC collaborative serializer · 0.9batching storage access · 0.9automatic field update · 0.9asynchronous APIs · 0.9microbenchmarking · 0.8case study · 0.8
YearPublicationVenuePosition
2026 SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC
abstract
As the gap between network and CPU speeds rapidly increases, the CPU-centric network stack proves inadequate due to excessive CPU and memory overheads. Though hardware-offloaded network stacks alleviate these issues, they suffer from limited flexibility in both control and data planes. It seems promising to offload network stacks to Smart-NICs to provide high flexibility. However, naive offloading leads to low throughput due to the inherent architectural limitations of widespread off-path SmartNICs. Even simple operations on staged network traffic would overwhelm the limited SmartNIC memory bandwidth. To this end, we design SmartNS, a SmartNIC-centric network stack with software transport programmability and line-rate packet processing capabilities. To tackle the limitations of SmartNIC-induced challenges, we propose a header-only offloading TX path and an unlimited-working-set in-cache processing RX path to minimize memory traffic to fit the wimpy SmartNIC memory bandwidth. To fully utilize the SmartNIC computing resources, we propose a programmable offloading engine to enable cloud providers to offload customized tasks along with the network stack processing. We prototype SmartNS using the widespread Nvidia BlueField-3 SmartNIC, and implement RoCEv2 and Solar transport protocols by leveraging SmartNS's software programmability. SmartNS achieves 2.2× higher throughput than the microkernel-based baseline in block storage disaggregation and 1.3× higher throughput than the hardware-offloaded baseline in KVCache transfer.
Xuzheng Chen, Jie Zhang 0081, Baolin Zhu, Xueying Zhu, Zhongqing Chen, Lingjun Zhu, Yin Zhang 0006, Yuanchao Shu, Peng Cheng 0001, Zeke Wang
EuroSys1
2026 Janus: Enabling Expressive and Efficient ACLs in High-speed RDMA Clouds
Ziteng Chen, Menghao Zhang 0001, Jiahao Cao 0001, Xuzheng Chen, Qiyang Peng
NDSS4
2025 RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs
abstract
The emerging microservice/serverless-based cloud programming paradigm and the rising networking speeds leave the RPC stack as the predominant data center tax. Domain-specific hardware acceleration holds the potential to disentangle the overhead and save host CPU cycles. However, state-of-the-art RPC accelerators integrate RPC logic into the CPU or use specialized low-latency interconnects, hardly adopted in commodity servers. To this end, we design and implement RpcNIC, a software-hardware co-designed SmartNIC that enables efficient RPC layer offloading and reconfigurable RPC kernel offloading. RpcNIC connects to the server through the most widely used PCIe interconnect. To grapple with the ramifications of PCIe-induced challenges, RpcNIC introduces three techniques: (a) a target-aware deserializer that effectively batches cross-PCIe writes on the SmartNIC’s SRAM using compacted hardware data structures; (b) a memory-affinity CPU-SmartNIC collaborative serializer, which trades additional host memory copies for slow cross PCIe-transfers; (c) an automatic field update technique that transparently codifies the schema based on dynamic reconfigure RPC kernels to minimize superfluous PCIe traversals. We prototype RpcNIC using the Xilinx U280 FPGA card. On HyperProtoBench, RpcNIC achieves an average of 2.3 × lower RPC layer processing time than a comparable RPC accelerator baseline and demonstrates 2.6 × achievable throughput improvement in the end-to-end cloud workload.
Jie Zhang 0081, Hongjing Huang, Xuzheng Chen, Xiang Li 0205, Jieru Zhao, Ming Liu 0027, Zeke Wang
HPCA3
2025 CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access
abstract
With the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems cost-effectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPU-managed SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to$\mathbf{1.84}\times, \mathbf{1.5}\times$, and$\mathbf{1.84}\times \mathbf{faster}$compared to the existing state-of-the-art GPU systems, while keeping high programmability.
Ziyu Song, Jie Zhang 0081, Jie Sun 0017, Mo Sun 0001, Zihan Yang 0004, Xuzheng Chen, Fei Wu 0001, Huajin Tang, Zeke Wang
ICDE7
2025 SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine
Hongjing Huang, Jie Zhang 0081, Xuzheng Chen, Ziyu Song, Jiajun Qin, Zeke Wang
USENIX ATC3
2024 DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications
abstract
Modern datacenter applications are increasingly being built using a microservices architecture. These microservices communicate with each other using datacenter RPCs. RPC's pass by value semantics incur redundant data movement along the network, especially for data-intensive applications. Naively introducing a shared global address space to datacenter RPC does not work as it would couple microservices and require microservices to handle data consistency, significantly complicating the development and deployment of applications. Fortunately, the modern datacenter is embracing disaggregated memory (DM). In a DM-enabled datacenter, servers running the microservices can be all connected to one global disaggregated memory pool, thus the pass by value semantics can be replaced by pass by reference. However, prior work on DM requires complicated synchronization primitives to share data across physical machines, so naively adopting them to datacenter RPC would harm microservices' agility and modularity. To this end, we present DmRPC, a DM-aware datacenter RPC for data-intensive datacenter applications to our knowledge. First, DmRPC introduces a DM-aware shared global address space to provide the semantics of pass by reference to datacenter RPC, thus alleviating the redundant data movement issue. Second, DmRPC adopts a copy-on-write mechanism to avoid complicating application logic to handle data consistency while guaranteeing high performance. We have applied DmRPC to two different implementations of DM, one is network-based (DmRPC-net) while the other is CXL-based (DmRPC-CXL). Our evaluations on synthetic 7-tier microservices workloads show that DmRPC-net (or DmRPC-CXL) achieves 4.2× (or 8.3×) higher throughput and achieves 1.1 × (or 1.7 ×) lower average latency than that of the baseline, respectively. On a widely used microservice benchmark DeathStarBench, DmRPC-net can achieve 3.1 × higher throughput and 2.5 × lower average latency than the baseline.
Jie Zhang 0081, Xuzheng Chen, Yin Zhang 0006, Zeke Wang
ICDE2
2024 Demystifying Datapath Accelerator Enhanced Off-path SmartNIC
abstract
Network speeds grow quickly in the modern cloud, so SmartNICs are introduced to offload network processing tasks, even application logic. However, typical multicore SmartNICs such as BlueFiled-2 are only capable of processing control-plane tasks with their embedded processors that have limited memory bandwidth and computing power. On the other hand, cloud applications evolve rapidly, such that a limited number of fixed hardware engines in a SmartNIC cannot satisfy the requirements of cloud applications. Therefore, SmartNIC programmers call for a programmable datapath accelerator (DPA) to process network traffic at line rate. However, no existing work has unveiled the performance characteristics of the existing DPA. To this end, we present the first architectural characterization of the latest DPA-enhanced BlueFiled-3 (BF3) SmartNIC. Our evaluation results indicate that BF3's DPA is significantly wimpier than the off-path Arm processor and the host CPU. However, we still identify that DPA has three unique architectural characteristics that unleash the performance potential of DPA. Specifically, we demonstrate how to take advantage of DPA's three architectural characteristics regarding computing, networking, and memory subsystems. Then we propose three important guidelines for programmers to fully unleash the potential of DPA. To demonstrate the effectiveness of our approach, we conduct detailed case studies regarding each guideline. Our case study on key-value aggregation achieves up to$4.3 \times$higher throughput by using our guidelines to optimize memory combinations.
Xuzheng Chen, Jie Zhang 0081, Lingjun Zhu, Yin Zhang 0006, Ming Liu 0027, Zeke Wang
ICNP1