Yongchen Pan

dblp:379/3745 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Tlaloc: A Generic Multipath Load Balancing for RoCE
Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
INFOCOM3
2026 LST-Sim: An Efficient Simulation Platform for Large-Scale Model Training
Siwei Ji, Bo Lei 0002, Yongchen Pan, Xiaoting Ma
SIGCOMM5
2026 Weir: Scalable RDMA With Delay-Based RNIC Cache Control Software Middleware for Data Center Networks
abstract
Remote Direct Memory Access (RDMA) is widely used in distributed services in Data Center Networks (DCNs) due to its high performance. As DCNs expand in scale, RDMA faces scalability issues. The reason is that the high concurrency Queue Pairs (QPs) lead to cache misses on RDMA Network Interface Card (RNIC) and frequent evictions, and the behaviour of fetching the cache via PCIe leads to performance degradation of RDMA. In this paper, we model the behaviour of Work Queue Element (WQE) on RNIC as a producer-consumer model and investigate that the root cause of WQE cache misses is the mismatch between the production rate of the CPU and the consumption rate of the RNIC. We design Weir from the perspective of WQE cache control to avoid cache misses and improve throughput under high concurrent QPs. Weir determines the cache occupancy on the RNIC by monitoring the number of active QPs and the increase/decrease in the life cycle of WQEs, and calculates the production rate and pacing by credit. The implementation of Weir exhibits minimal CPU overhead. Evaluation results show that Weir can maintain 97Gbps throughput without degradation even with up to 16K concurrent QPs, and effectively reduces various observable cache misses by$5\times $to$10\times $compared to commercial RNICs. Additionally, experiments show that Weir has better connection scalability than XRC and DCT.
Jiao Zhang 0002, Yongchen Pan, Dexuan Liao, Huimin Luo, Tao Huang 0005, Haipeng Yao
IEEE Trans. Netw.3
2025 Valve: Scalable RDMA with Gap-based Cache Control Middleware for Data Center Networks
abstract
Remote Direct Memory Access (RDMA) has become a cornerstone technology in Data Center Networks (DCNs). However, DCNs have expanded substantially, leading to severe connection scalability issues for RDMA. The critical reason behind these issues stems from frequent cache misses on RDMA NICs (RNICs) when handling numerous Queue Pair (QP) connections. Cache misses require time-consuming retrievals from the host via PCIe, resulting in a degradation in RDMA performance. Existing software solutions primarily aim to alleviate QP Context (QPC) cache pressure, while hardware solutions incur prohibitive costs. In this paper, we identify Work Queue Element (WQE) cache, rather than QPC cache, as the fundamental bottleneck limiting connection scalability. Hence, we propose a software middleware, Valve, designed to mitigate cache misses through WQE cache control. Valve regulates WQE posting to control WQE cache by monitoring RNIC cache usage and adaptively adjusting the gap of WQE posting. Valve boasts ease of deployment, requiring the addition of approximately 1000 lines of code, and incurs low CPU overhead. Valve maintains peak performance of RNIC throughput regardless of the number of QPs and considerably reduces observable cache misses (such as ICM, MTT, and MPT cache misses) by 2.8× to 3.1× compared to XRC and DCT.
Jiao Zhang 0002, Dexuan Liao, Yongchen Pan, Tao Huang 0005
ICNP5
2025 Weir: Delay-based RNIC Cache Control Software Middleware for Scalable RDMA Networks
Yongchen Pan, Jiao Zhang 0002, Zirui Wan, Baohong Lin, Junliang Wang, Huimin Luo
INFOCOM1
2025 SeqBalance: Congestion-Aware Load Balancing With No Reordering in Data Center Networks
abstract
With the rapid development of the Internet of Things (IoT), an increasing amount of sensor data generated by IoT applications has been transferred to data center networks for storage and data analysis. Remote Direct Memory Access (RDMA) is widely used in data center networks because of its high performance. However, due to the characteristics of RDMA’s retransmission strategy, current load balancing schemes for data center networks are unsuitable for RDMA. In this paper, we propose SeqBalance, a load balancing framework designed for RDMA. SeqBalance implements fine-grained load balancing for RDMA through a reasonable design and does not cause reordering problems. SeqBalance detects link congestion at the switch by sensing ECN signals and link utilization, and guides routing decisions accordingly. SeqBalance’s designs are all based on existing commercial RNICs and commercial programmable switches, so they are compatible with existing data center networks. We have implemented SeqBalance Shaper for fine-grained sub-flow splitting in Mellanox CX-6 RNIC and implemented routing decisions in Intel Tofino P4 programmable switch. The results of hardware testbed experiments and large-scale simulations show that compared with existing load balancing schemes, SeqBalance improves 24.7% and 15.9% on average FCT and 99th-percentile FCT.
Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
IEEE Internet Things J.4
2025 RoCELet: Host-Based Flowlet Load Balancing for RoCE
abstract
Remote Direct Memory Access (RDMA) is becoming a popular high-speed networking technology. It uses kernel bypass and zero copy to achieve high throughput and low latency with little CPU overhead. However, standard RoCE transmission uses Equal Cost Multipath (ECMP) for load balancing, which can result in lower transmission performance due to hash conflicts. Meanwhile, it has been verified that, unlike TCP, the unique retransmission mode and flow characteristics of RoCE make previous load balancing algorithms not well applied to RoCE. In this paper, we introduce RoCELet, a load balancing algorithm for RoCE. It achieves fine-grained RoCE load balancing by actively generating flowlets, effectively utilizing the rich end-to-end paths in the data center. We implement a prototype based on DPDK and evaluate it through small-scale testbed experiments and large-scale simulations. Our results show that compared to state-of-the-art load balancing algorithms, RoCELet optimizes 48.2% and 16.4% in average FCT and$99^{th}$-ile FCT, respectively.
Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Jiafeng Jiang, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
IEEE Trans. Netw.5
2025 RHCC: Revisiting Intra-Host Congestion Control in RDMA Networks
abstract
RDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with CPU-to-memory traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to revisit the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines CPU-to-memory traffic congestion avoidance with sub-RTT granularity and proactive RNIC traffic adjustment. RHCC ensures fast congestion avoidance and can work with different inter-host congestion control methods. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2$\times$and 1.4$\times$, respectively.
Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Yongchen Pan, Tao Huang 0005
IEEE Trans. Netw.6
2024 Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE Clusters
abstract
RoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary.
Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Zicheng Wang 0004, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
APNet13
2024 An Integrated Method for Fast Imaging and Detection of Lightweight Intelligent Ship Targets
abstract
Ship target detection based on SAR images is an important means of marine observation. Traditional target detection requires the most processing time to image the SAR echo. Considering the sparse distribution of ship targets in wide-swath marine SAR images, imaging and detecting processes on non-target regions seriously reduce efficiency. This paper proposes an integrated framework to improve marine SAR imaging detection efficiency by adding two steps of selection for target areas. Firstly, an RC-TextCNN network is designed to select target areas on azimuth direction from SAR echo one-dimensional compression data. After imaging selected areas, a dynamic quantization and threshold segmentation method is used to further remove non-target areas. Finally, suspected target areas are introduced into the pruned yolov7 model for final target detection. This workflow significantly minimizes computational and time costs. The experiment on Gaofen3 data shows that the speed of the process is increased by three times while detection accuracy is at 90%.
Can Su, Yongchen Pan, Wei Yang 0004, Hongcheng Zeng 0001
IGARSS2
2024 R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic System
abstract
RoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related.
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Haohan Xu, Dongyang Song, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
SIGCOMM17
2024 Blaze: Delay-Aware Cloud-Edge Collaborative Service Function Chain Deployment with Network Calculus
abstract
With the rapid development of Internet of the Things (IoT) technology, IoT services have higher and higher requirements for latency. In the IoT environment, virtual network functions (VNFs) are deployed on general-purpose hardware and are sequentially connected to form service function chain (SFC) to provide network services for IoT devices. However, the high latency of the link between the cloud center and the edge nodes and the resource capacity limitation of the edge nodes pose challenges to the deployment of SFCs in IoT devices. In this paper, we study the cloud-edge collaborative SFC deployment problem. We applied the network calculus theory to the cloud-edge collaborative SFC deployment for the first time, aiming to provide the end-to-end delay guarantee for the deployed SFC. We model the SFC deployment problem as Mixed Integer Nonlinear Programming (MINLP). Then we propose a heuristic algorithm (Blaze) to solve this problem. Blaze is proven to complete the deployment of SFCs in polynomial time. Finally, the algorithm is evaluated by experimental simulation. The experimental results show that compared with the existing state-of-the-art corresponding algorithms, the proposed algorithm achieves better performance in terms of the number of VNFs deployed in the cloud, resource consumption of edge nodes, and SFC request acceptance rate.
Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
WCNC3