EDBT 2026 Demo / reviewers in the wild / expert
Hao Wang 0231
dblp:181/2812-231
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0001-7846-9840ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training
Huifeng Xing, Hao Wang 0231, Yinfan Hu, Xin Ai 0008, Yang Chen 0001, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
INFOCOM | 2 |
| 2026 | Scale-up PIFO: Interleaving Multiple Priority Queues for High Speed Programmable SchedulingabstractPush-In First-Out (PIFO) offers a unified abstraction for rapidly deploying diverse scheduling algorithms on the same hardware. As SerDes-lane aggregation pushes port rates to 1.6 Tbps, the perpacket processing budget is at a sub-nanosecond scale, making single-queue PIFO designs fail to keep up. Mirroring lane aggregation, we advocate interleaving multiple PIFO queues. However, simple round-robin parallelization introduces substantial scheduling error, and in the worst case, it can grow to the order of the buffer size. Shili Chen, Ruyi Yao, Zhiyu Zhang 0012, Hao Wang 0231, Deli Huang, Yibo Fan, Yang Xu 0010 |
SIGCOMM | 7 |
| 2026 | MetaFlex: A Flexible Architecture for Efficient Packet Scheduling and Memory AllocationabstractPacket schedulers are essential for managing packet transmission order in high-speed networks, where scheduling metadata must be processed at line rate under bursty traffic conditions. In such systems, packet scheduling and traffic management operate on compact packet descriptors rather than on full packets. FIFO-based schedulers are attractive for their simplicity and throughput, but implementations that statically partition descriptor memory across queues must provision for worst-case occupancy, leading to inefficient memory utilization. This paper presents MetaFlex, a scalable architecture for efficient implementation of calendar-queue–based packet scheduling using shared descriptor memory. MetaFlex dynamically allocates descriptor storage to scheduling queues on demand, allowing memory to be effectively shared across a large number of rank bins while maintaining constant-time enqueue and dequeue operations. Rather than introducing a new scheduling algorithm, MetaFlex focuses on the architectural realization of shared-memory descriptor queues that scale to large queue counts with predictable timing and modest hardware cost. We evaluate MetaFlex using NS2 simulations of weighted fair scheduling and a hardware prototype implemented in VHDL on an AMD/Xilinx Alveo U250 FPGA board. Simulation results show that MetaFlex achieves comparable or lower packet loss than fixed-memory calendar queues under identical descriptor-memory budgets, while using substantially less descriptor storage under typical traffic conditions. The FPGA prototype operates at 322 MHz, sustains 100 Gb/s line rate for packets larger than 370 bytes, and uses less than 1% of logic resources and less than 5% of on-chip memory, demonstrating the practicality of MetaFlex for high-speed hardware datapaths. Anthony Dalleggio, Peixuan Gao, Yongbo Gao, Hao Wang 0231, Yang Xu 0010, H. Jonathan Chao |
IEEE Trans. Netw. | 4 |
| 2026 | AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource PoolingabstractThe increasing popularity of large models and datasets has highlighted the significance of distributed training networks. As gradient synchronization generates substantial traffic, in-network aggregation (INA) has emerged as a solution to offload aggregation onto the switch, alleviating network congestion and accelerating distributed training. However, the limited memory capacity of the INA switch becomes a potential bottleneck as computation shifts into the network, especially in multi-tenant scenarios. To address this bottleneck and enhance network throughput, we propose the Aggregation with Innetwork Resource Pooling (AIRP) framework. Unlike existing approaches that optimize individual switches in a localized manner, AIRP takes a holistic view and efficiently pools switch memory resources across the entire network, allocating them to multiple tenants. Evaluation using the ns-3 simulator and P4 testbed demonstrates that AIRP can accelerate the training of various models, including computer vision and language models. The experimental results show that AIRP outperforms existing INA approaches by up to 7 times in terms of network throughput in multi-tenant scenarios, while also achieving great flexibility and efficiency in deployment. Huifeng Xing, Hao Wang 0231, Yang Chen 0001, Yinfan Hu, Xuandong Liu, Zijian Li 0003, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
IEEE Trans. Netw. | 2 |
| 2025 | Enhancing In-network Aggregation with Adaptive Gradient Quantization for Multi-tenant LearningabstractWith the increasing popularity of distributed training applications, the growth in network traffic has become an impediment to the communication among worker nodes in the system. In-network aggregation (INA) has emerged as a solution to improve communication efficiency by offloading gradient aggregation to switches. However, in multi-tenant scenarios, INA switch memory capacity has been identified as a main bottleneck, leading to reduced network throughput and slower training processes. To address this, we propose Adaptive Gradient Quantization (AGQ) on the switch. AGQ reduces the quantization bit-width of gradients, allowing for storage of more gradients within the limited switch memory while maintaining training accuracy. Compared to quantization on hosts, AGQ can swiftly adapt to the available memory on switches and offers an improved balance between minimizing precision loss and enhancing training throughput. We implement AGQ on a P4 switch testbed, and experimental results demonstrate that enabling AGQ can achieve an up to 100% increase in training throughput without explicit drop of training accuracy compared with existing INA solutions like ATP and host-based quantization methods like THC. Huifeng Xing, Yinfan Hu, Hao Wang 0231, Yang Chen 0001, Sen Liu 0002, Yang Xu 0010 |
ICDCS | 3 |
| 2024 | Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model TrainingabstractParameter Server (PS) and Ring-AllReduce (RAR) are two widely utilized synchronization architectures in multiworker Deep Learning (DL), also referred to as Distributed Deep Learning (DDL). However, PS encounters challenges with the “incast” issue, while RAR struggles with problems caused by the long dependency chain. The emerging In-network Aggregation (INA) has been proposed to integrate with PS to mitigate its incast issue. However, such PS-based INA has poor incremental deployment abilities as it requires replacing all the switches to show significant performance improvement, which is not costeffective. In this study, we present the incorporation of INA capabilities into RAR, called RAR with In-Network Aggregation (Rina), to tackle both the problems above. Rina features its agent-worker mechanism. When an INA-capable ToR switch is deployed, all workers in this rack run as one abstracted worker with the help of the agent, resulting in both excellent incremental deployment capabilities and better throughput. We conducted extensive testbed and simulation evaluations to substantiate the throughput advantages of Rina over existing DDL training synchronization structures. Compared with the state-of-the-art PS-based INA methods ATP, Rina can achieve more than$\mathbf{5 0 \%}$throughput with the same hardware cost. Xuandong Liu, Minglin Li, Yinfan Hu, Huifeng Xing, Hao Wang 0231, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
ICNP | 7 |
| 2024 | TCAMVisor: High-throughput TCAM Virtualization for Multi-tenant Software Defined NetworkingabstractSoftware Defined Networking (SDN) provides users with a unified abstraction of physical networks. To meet the demands of modern data centers, many works have focused on designing network virtualization hypervisors that support multi-tenant SDN. Ternary Content Addressable Memory (TCAM) is widely used in SDN switches for rule storage. While it has extremely high lookup throughput, it also features drawbacks such as small capacity and slow update speed. Faced with multi-tenant scenarios, its limitations are even more pronounced. Existing hypervisors lack consideration for TCAM isolation, leading to slower TCAM updates and the mutual impact of requests from different tenants. Consequently, they fail to provide guaranteed performance to tenants. To solve these problems, we propose TCAMVisor, which further isolates TCAM resources based on traditional SDN hypervisors. Specifically, TCAMVisor provides better allocation mechanisms for TCAM entry and control bandwidth, ensuring inter-tenant isolation while improving resource utilization. Additionally, TCAMVisor improves the update speed of TCAM by delicately placing tenant rules. To the best of our knowledge, TCAMVisor is the first work to effectively achieve tenant isolation in TCAM, with an average throughput improvement of 5.5 times compared to FlowVisor. Ruoshi Sun, Ruyi Yao, Hao Wang 0231, Yiren Zhou, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 4 |
| 2024 | vPIFO: Virtualized Packet Scheduler for Programmable Hierarchical Scheduling in High-Speed NetworksabstractProgrammable packet scheduling enables the integration of scheduling algorithms into switches without the need for hardware redesign. The Push-In First-Out (PIFO) queue facilitates a programmable packet scheduler, supporting a single scheduling algorithm flexibly. However, hierarchical scheduling required in Multi-Tenant Data Centers (MTDCs) remains non-programmable. Dynamic and diverse hierarchical scheduling algorithms necessitate alterations in both the number of PIFO queues and their connection topology, posing a significant challenge to support them on fixed hardware. Zhiyu Zhang 0012, Shili Chen, Ruyi Yao, Ruoshi Sun, Hao Wang 0231, Gaojian Fang, Yibo Fan, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
SIGCOMM | 6 |