EDBT 2026 Demo / reviewers in the wild / expert
Xuzheng Chen
dblp:367/9362
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0009-7432-5485ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Cloud and datacenter computing · 33% GPUs and heterogeneous computing · 15% Hardware accelerators and domain-specific architectures · 14% | |
| Computer networks
2 papers |
Datacenter networks · 53% Software-defined and programmable networks · 47% | |
| Network and information security
1 paper |
Authentication and access control · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › network accelerator
SmartNIC |
1.6 | 2 | 2025 | RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025 Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024 |
Software-defined and programmable networks
programmable data plane |
1.0 | 1 | 2026 | SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC · EuroSys 2026 |
Authentication and access control
access control |
1.0 | 1 | 2026 | Janus: Enabling Expressive and Efficient ACLs in High-speed RDMA Clouds · NDSS 2026 |
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload |
1.0 | 1 | 2026 | SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNIC · EuroSys 2026 |
Datacenter networks › RDMA
RDMA congestion control |
0.9 | 1 | 2025 | SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine · USENIX ATC 2025 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems |
0.9 | 1 | 2025 | CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025 |
Cloud and datacenter computing
datacenter network |
0.9 | 1 | 2025 | RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025 |
GPUs and heterogeneous computing
GPU storage access |
0.9 | 1 | 2025 | CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025 |
Distributed systems › remote procedure call
RPC offloading |
0.9 | 1 | 2025 | RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025 |
Performance modeling and evaluation › workload characterization
architectural characterization |
0.8 | 1 | 2024 | Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024 |
Cloud and datacenter computing
datacenter RPC |
0.8 | 1 | 2024 | DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024 |
Distributed systems
remote procedure call |
0.8 | 1 | 2024 | DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024 |
Performance modeling and evaluation
workload characterization |
0.8 | 1 | 2024 | Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024 |
Datacenter networks
RDMA |
0.3 | 1 | 2025 | SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine · USENIX ATC 2025 |
Storage systems › flash and SSD › solid-state drive
NVMe SSD |
0.3 | 1 | 2025 | CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025 |
Interconnection networks and networks-on-chip › high-speed interconnect
PCIe interconnect |
0.3 | 1 | 2025 | RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs · HPCA 2025 |
Memory systems › memory disaggregation
CXL memory disaggregation |
0.2 | 1 | 2024 | DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024 |
Memory systems
memory disaggregation |
0.2 | 1 | 2024 | DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications · ICDE 2024 |
Cloud and datacenter computing › computation offloading
network function offloading |
0.2 | 1 | 2024 | Demystifying Datapath Accelerator Enhanced Off-path SmartNIC · ICNP 2024 |
Methods — techniques the papers use, named apart from their topics
in-cache processing · 2.0header-only offloading · 2.0RDMA · 2.0target-aware deserializer · 0.9software-programmable congestion control · 0.9memory-affinity CPU-SmartNIC collaborative serializer · 0.9batching storage access · 0.9automatic field update · 0.9asynchronous APIs · 0.9microbenchmarking · 0.8case study · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICabstractAs the gap between network and CPU speeds rapidly increases, the CPU-centric network stack proves inadequate due to excessive CPU and memory overheads. Though hardware-offloaded network stacks alleviate these issues, they suffer from limited flexibility in both control and data planes. It seems promising to offload network stacks to Smart-NICs to provide high flexibility. However, naive offloading leads to low throughput due to the inherent architectural limitations of widespread off-path SmartNICs. Even simple operations on staged network traffic would overwhelm the limited SmartNIC memory bandwidth. To this end, we design SmartNS, a SmartNIC-centric network stack with software transport programmability and line-rate packet processing capabilities. To tackle the limitations of SmartNIC-induced challenges, we propose a header-only offloading TX path and an unlimited-working-set in-cache processing RX path to minimize memory traffic to fit the wimpy SmartNIC memory bandwidth. To fully utilize the SmartNIC computing resources, we propose a programmable offloading engine to enable cloud providers to offload customized tasks along with the network stack processing. We prototype SmartNS using the widespread Nvidia BlueField-3 SmartNIC, and implement RoCEv2 and Solar transport protocols by leveraging SmartNS's software programmability. SmartNS achieves 2.2× higher throughput than the microkernel-based baseline in block storage disaggregation and 1.3× higher throughput than the hardware-offloaded baseline in KVCache transfer. Xuzheng Chen, Jie Zhang 0081, Baolin Zhu, Xueying Zhu, Zhongqing Chen, Lingjun Zhu, Yin Zhang 0006, Yuanchao Shu, Peng Cheng 0001, Zeke Wang |
EuroSys | 1 |
| 2026 | Janus: Enabling Expressive and Efficient ACLs in High-speed RDMA Clouds
Ziteng Chen, Menghao Zhang 0001, Jiahao Cao 0001, Xuzheng Chen, Qiyang Peng |
NDSS | 4 |
| 2025 | RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsabstractThe emerging microservice/serverless-based cloud programming paradigm and the rising networking speeds leave the RPC stack as the predominant data center tax. Domain-specific hardware acceleration holds the potential to disentangle the overhead and save host CPU cycles. However, state-of-the-art RPC accelerators integrate RPC logic into the CPU or use specialized low-latency interconnects, hardly adopted in commodity servers. To this end, we design and implement RpcNIC, a software-hardware co-designed SmartNIC that enables efficient RPC layer offloading and reconfigurable RPC kernel offloading. RpcNIC connects to the server through the most widely used PCIe interconnect. To grapple with the ramifications of PCIe-induced challenges, RpcNIC introduces three techniques: (a) a target-aware deserializer that effectively batches cross-PCIe writes on the SmartNIC’s SRAM using compacted hardware data structures; (b) a memory-affinity CPU-SmartNIC collaborative serializer, which trades additional host memory copies for slow cross PCIe-transfers; (c) an automatic field update technique that transparently codifies the schema based on dynamic reconfigure RPC kernels to minimize superfluous PCIe traversals. We prototype RpcNIC using the Xilinx U280 FPGA card. On HyperProtoBench, RpcNIC achieves an average of 2.3 × lower RPC layer processing time than a comparable RPC accelerator baseline and demonstrates 2.6 × achievable throughput improvement in the end-to-end cloud workload. Jie Zhang 0081, Hongjing Huang, Xuzheng Chen, Xiang Li 0205, Jieru Zhao, Ming Liu 0027, Zeke Wang |
HPCA | 3 |
| 2025 | CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessabstractWith the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems cost-effectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPU-managed SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to$\mathbf{1.84}\times, \mathbf{1.5}\times$, and$\mathbf{1.84}\times \mathbf{faster}$compared to the existing state-of-the-art GPU systems, while keeping high programmability. Ziyu Song, Jie Zhang 0081, Jie Sun 0017, Mo Sun 0001, Zihan Yang 0004, Xuzheng Chen, Fei Wu 0001, Huajin Tang, Zeke Wang |
ICDE | 7 |
| 2025 | SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine
Hongjing Huang, Jie Zhang 0081, Xuzheng Chen, Ziyu Song, Jiajun Qin, Zeke Wang |
USENIX ATC | 3 |
| 2024 | DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive ApplicationsabstractModern datacenter applications are increasingly being built using a microservices architecture. These microservices communicate with each other using datacenter RPCs. RPC's pass by value semantics incur redundant data movement along the network, especially for data-intensive applications. Naively introducing a shared global address space to datacenter RPC does not work as it would couple microservices and require microservices to handle data consistency, significantly complicating the development and deployment of applications. Fortunately, the modern datacenter is embracing disaggregated memory (DM). In a DM-enabled datacenter, servers running the microservices can be all connected to one global disaggregated memory pool, thus the pass by value semantics can be replaced by pass by reference. However, prior work on DM requires complicated synchronization primitives to share data across physical machines, so naively adopting them to datacenter RPC would harm microservices' agility and modularity. To this end, we present DmRPC, a DM-aware datacenter RPC for data-intensive datacenter applications to our knowledge. First, DmRPC introduces a DM-aware shared global address space to provide the semantics of pass by reference to datacenter RPC, thus alleviating the redundant data movement issue. Second, DmRPC adopts a copy-on-write mechanism to avoid complicating application logic to handle data consistency while guaranteeing high performance. We have applied DmRPC to two different implementations of DM, one is network-based (DmRPC-net) while the other is CXL-based (DmRPC-CXL). Our evaluations on synthetic 7-tier microservices workloads show that DmRPC-net (or DmRPC-CXL) achieves 4.2× (or 8.3×) higher throughput and achieves 1.1 × (or 1.7 ×) lower average latency than that of the baseline, respectively. On a widely used microservice benchmark DeathStarBench, DmRPC-net can achieve 3.1 × higher throughput and 2.5 × lower average latency than the baseline. Jie Zhang 0081, Xuzheng Chen, Yin Zhang 0006, Zeke Wang |
ICDE | 2 |
| 2024 | Demystifying Datapath Accelerator Enhanced Off-path SmartNICabstractNetwork speeds grow quickly in the modern cloud, so SmartNICs are introduced to offload network processing tasks, even application logic. However, typical multicore SmartNICs such as BlueFiled-2 are only capable of processing control-plane tasks with their embedded processors that have limited memory bandwidth and computing power. On the other hand, cloud applications evolve rapidly, such that a limited number of fixed hardware engines in a SmartNIC cannot satisfy the requirements of cloud applications. Therefore, SmartNIC programmers call for a programmable datapath accelerator (DPA) to process network traffic at line rate. However, no existing work has unveiled the performance characteristics of the existing DPA. To this end, we present the first architectural characterization of the latest DPA-enhanced BlueFiled-3 (BF3) SmartNIC. Our evaluation results indicate that BF3's DPA is significantly wimpier than the off-path Arm processor and the host CPU. However, we still identify that DPA has three unique architectural characteristics that unleash the performance potential of DPA. Specifically, we demonstrate how to take advantage of DPA's three architectural characteristics regarding computing, networking, and memory subsystems. Then we propose three important guidelines for programmers to fully unleash the potential of DPA. To demonstrate the effectiveness of our approach, we conduct detailed case studies regarding each guideline. Our case study on key-value aggregation achieves up to$4.3 \times$higher throughput by using our guidelines to optimize memory combinations. Xuzheng Chen, Jie Zhang 0081, Lingjun Zhu, Yin Zhang 0006, Ming Liu 0027, Zeke Wang |
ICNP | 1 |