Ke Wu 0003

dblp:69/6116-3 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-0625-3004ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 HNCC: Host-Network Collaborative Congestion Control for RDMA
abstract
Driven by the rapidly growing demands of modern applications, RDMA technology has been widely adopted in cluster networks, including high-performance computing (HPC) systems and datacenters. However, the high-speed, kernel-bypass nature of RDMA presents significant challenges for congestion control (CC). On one hand, since most RDMA operations are offloaded to the RDMA network interface card (RNIC), software-based CC struggles to regulate packet transmission effectively. On the other hand, RDMA workloads are dominated by short messages, making it difficult for both reactive and proactive CC protocols to manage congestion in a timely and effective manner, posing risks to network application QoS. Despite these challenges, proactive and reactive CC protocols each have distinct strengths, and combining them offers a promising research direction. This paper proposes an endpoint-network collaborative congestion control protocol (HNCC) that integrates proactive endpoint congestion control with reactive network monitoring and adjustment to achieve precise and rapid congestion management. HNCC introduces a novel reservation mechanism embedded within the RDMA protocol to prevent endpoint congestion while efficiently mitigating in-network congestion by dynamically adjusting the data injection rate based on network trend analysis. HNCC is offloaded to the RNIC, ensuring that its traffic scheduling is timely and effective. Extensive load testing across various scales shows that HNCC achieves performance comparable to state-of-the-art CC protocols in datacenter workloads and reduces the average flow completion time (FCT) by 5 − 56% under HPC workloads.
Yuang Yang, Dezun Dong, Ke Wu 0003, Yunyang Xu
IWQoS3
2025 A lightweight RDMA connection protocol based on post-hoc confirmation
Ke Wu 0003, Dezun Dong, Weixia Xu 0001
J. Parallel Distributed Comput.1
2024 DRLAR: A deep reinforcement learning-based adaptive routing framework for network-on-chips
Ke Wu 0003, Cunlu Li, Dezun Dong
Comput. Networks4
2024 COER: A Network Interface Offloading Architecture for RDMA and Congestion Control Protocol Codesign
abstract
RDMA (Remote Direct Memory Access) networks require efficient congestion control to maintain their high throughput and low latency characteristics. However, congestion control protocols deployed at the software layer suffer from slow response times due to the communication overhead between host hardware and software. This limitation has hindered their ability to meet the demands of high-speed networks and applications. Harnessing the capabilities of rapidly advancing Network Interface Cards (NICs) can drive progress in congestion control. Some simple congestion control protocols have been offloaded to RDMA NICs to enable faster detection and processing of congestion. However, offloading congestion control to the RDMA NIC faces a significant challenge in integrating the RDMA transport protocol with advanced congestion control protocols that involve complex mechanisms. We have observed that reservation-based proactive congestion control protocols share strong similarities with RDMA transport protocols, allowing them to integrate seamlessly and combine the functionalities of the transport layer and network layer. In this article, we present COER, an RDMA NIC architecture that leverages the functional components of RDMA to perform reservations and completes the scheduling of congestion control during the scheduling process of the RDMA protocol. COER facilitates the streamlined development of offload strategies for congestion control techniques —specifically, proactive congestion control —on RDMA NICs. We use COER to design offloading schemes for 11 congestion control protocols, which we implement and evaluate using a network emulator with a cycle-accurate RDMA NIC model that can load Message Passing Interface (MPI) programs. The evaluation results demonstrate that the architecture of COER does not compromise the original characteristics of the congestion control protocols. Compared with a layered protocol stack approach, COER enables the performance of RDMA networks to reach new heights.
Ke Wu 0003, Dezun Dong, Weixia Xu 0001
ACM Trans. Archit. Code Optim.1
2023 DFR: Dynamic-thresold Fault-tolerant Routing for Fat Tree
Binyan Lan, Ke Wu 0003, Dezun Dong
APNet3
2023 DFAR: Dynamic-threshold Fault-tolerant Adaptive Routing for Fat Tree Networks
abstract
The routing algorithm is important for the design of high-performance interconnection networks. With the increasing size of high-performance computing (HPC) systems, the possibility of network component failures increases simultaneously. Fault tolerance becomes a more critical consideration for the routing algorithms, as failed network devices, mostly network links, corrupt the regularity of the topology. However, existing routing algorithms focus on load balancing, ignoring that when the network is faulty. Similarly, current fault-tolerant routing algorithms ensure the correct functionality when failures exist without considering the more challenging post-failure load balancing. In this paper, we co-design the load balancing and fault tolerance for adaptive routing algorithms in fat tree networks and propose DFAR, Dynamic-threshold Fault-tolerant Adaptive Routing. DFAR prioritizes D-Mod-K deterministic output ports by adding thresholds to other available candidates. We adopt a simulated gradient descent algorithm to dynamically update the thresholds according to the network state changes. More state information other than local queue occupancies is used in the thresholds optimization. Experiments with synthetic load show that DFAR improves throughput by up to 25% for network with considerable failures. For realistic MPI workloads, DFAR also improves performance by up to 22%.
Binyan Lan, Dezun Dong, Ke Wu 0003
ICPADS4
2023 Roar: A Router Microarchitecture for In-network Allreduce
abstract
The allreduce operation is the most commonly used collective operation in distributed or parallel applications. It aggregates data collected from distributed hosts and broadcasts the aggregated result back to them. In-network computing can accelerate allreduce by offloading this operation into network devices. However, existing in-network solutions face the challenge of high throughput, performance of aggregating large message and producing repeatable results. In this work, we propose a simple and effective router microarchitecture for in-network allreduce, which uses an RDMA protocol to improve its throughput. We further discuss strategies to tackle the aforementioned challenges. Our approach not only shows advantages in comparison with the state-of-the-art in-network solutions, but also accelerates allreduce at a near-optimal level compared to host-based algorithms, as demonstrated through experiments.
Dezun Dong, Ke Wu 0003, Kai Lu 0001
ICS5
2022 Revisiting network congestion avoidance through adaptive packet-chaining reservation
Ke Wu 0003, Dezun Dong, Cunlu Li, Weixia Xu 0001
Comput. Networks1
2021 Evaluation of Topology-Aware All-Reduce Algorithm for Dragonfly Networks
Dezun Dong, Cunlu Li, Ke Wu 0003, Liquan Xiao
NPC4
2019 PPS: A Low-Latency and Low-Complexity Switching Architecture Based on Packet Prefetch and Arbitration Prediction
Ke Wu 0003, Dezun Dong
ICA3PP (1)2
2019 Network Congestion Avoidance through Packet-chaining Reservation
abstract
Endpoint congestion is a bottleneck in high-performance computing (HPC) networks and severely impacts system performance, especially for latency-sensitive applications. For long messages (or flows) whose duration is far larger than the round-trip time (RTT), endpoint congestion can be effectively mitigated by proactive or reactive counter-measures such that the injection rate of each source is dynamically controlled to a proper level. However, many HPC applications produce a hybrid traffic, a mix of short and long messages, and are dominated by short messages. Existing proactive congestion avoidance methods face the great challenge of scheduling the rapidly changing traffic pattern caused by these short messages. In this paper, we leverage the advantages of proactive and reactive congestion avoidance techniques and propose the Packet-chaining Reservation Protocol (PCRP) to make a dynamic balance between flows following proactive scheduling and packets subjected to reactive network conditions. We select the chaining packets as a flexible reservation granularity between the whole flow and one packet. We allow small flows to be speculatively transmitted without being discarded and give them higher priority over the entire network. Our PCRP can respond quickly to network conditions and effectively avoid the formation of endpoint congestion and reduce the average flow delay. We conduct extensive experiments to evaluate our PCRP and compare it with the state-of-the-art proactive reservation-based protocols, Speculative Reservation Protocol (SRP) and Bilateral Flow Reservation Protocol (BFRP). The simulation results demonstrate that in our design the flow latency can be reduced by 50.2% for hotspot traffic and 28.38% for uniform traffic.
Ke Wu 0003, Dezun Dong, Cunlu Li, Shan Huang 0002
ICPP1