Yinfan Hu

dblp:383/6840 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0005-1775-7258ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 76% Parallel and multicore computing · 16% Cloud and datacenter computing · 5%
Computer networks
2 papers
Software-defined and programmable networks · 57% Internet of things and sensor networks · 43%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed machine learning
distributed training
2.832026
AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource Pooling · IEEE Trans. Netw. 2026
3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training · INFOCOM 2026
Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model Training · ICNP 2024
Distributed systems › data aggregation
in-network aggregation
2.022026
AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource Pooling · IEEE Trans. Netw. 2026
3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training · INFOCOM 2026
Software-defined and programmable networks › programmable data plane
programmable switch
1.012026
AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource Pooling · IEEE Trans. Netw. 2026
Parallel and multicore computing › parallel computing › parallel machine learning
3d parallelism
1.012026
3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training · INFOCOM 2026
Internet of things and sensor networks › wireless sensor network
in-network aggregation
0.812024
Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model Training · ICNP 2024
Cloud and datacenter computing › resource management
multi-tenant resource management
0.312026
AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource Pooling · IEEE Trans. Netw. 2026
High-performance computing
collective communication
0.212024
Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model Training · ICNP 2024

Methods — techniques the papers use, named apart from their topics

in-network aggregation · 3.0testbed evaluation · 1.5simulation · 1.5agent-worker mechanism · 1.5programmable data planes · 1.0programmable data plane · 1.03d parallelism · 1.0
YearPublicationVenuePosition
2026 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training
Huifeng Xing, Hao Wang 0231, Yinfan Hu, Xin Ai 0008, Yang Chen 0001, Wanxin Shi, Sen Liu 0002, Yang Xu 0010
INFOCOM3
2026 AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource Pooling
abstract
The increasing popularity of large models and datasets has highlighted the significance of distributed training networks. As gradient synchronization generates substantial traffic, in-network aggregation (INA) has emerged as a solution to offload aggregation onto the switch, alleviating network congestion and accelerating distributed training. However, the limited memory capacity of the INA switch becomes a potential bottleneck as computation shifts into the network, especially in multi-tenant scenarios. To address this bottleneck and enhance network throughput, we propose the Aggregation with Innetwork Resource Pooling (AIRP) framework. Unlike existing approaches that optimize individual switches in a localized manner, AIRP takes a holistic view and efficiently pools switch memory resources across the entire network, allocating them to multiple tenants. Evaluation using the ns-3 simulator and P4 testbed demonstrates that AIRP can accelerate the training of various models, including computer vision and language models. The experimental results show that AIRP outperforms existing INA approaches by up to 7 times in terms of network throughput in multi-tenant scenarios, while also achieving great flexibility and efficiency in deployment.
Huifeng Xing, Hao Wang 0231, Yang Chen 0001, Yinfan Hu, Xuandong Liu, Zijian Li 0003, Wanxin Shi, Sen Liu 0002, Yang Xu 0010
IEEE Trans. Netw.4
2025 Enhancing In-network Aggregation with Adaptive Gradient Quantization for Multi-tenant Learning
abstract
With the increasing popularity of distributed training applications, the growth in network traffic has become an impediment to the communication among worker nodes in the system. In-network aggregation (INA) has emerged as a solution to improve communication efficiency by offloading gradient aggregation to switches. However, in multi-tenant scenarios, INA switch memory capacity has been identified as a main bottleneck, leading to reduced network throughput and slower training processes. To address this, we propose Adaptive Gradient Quantization (AGQ) on the switch. AGQ reduces the quantization bit-width of gradients, allowing for storage of more gradients within the limited switch memory while maintaining training accuracy. Compared to quantization on hosts, AGQ can swiftly adapt to the available memory on switches and offers an improved balance between minimizing precision loss and enhancing training throughput. We implement AGQ on a P4 switch testbed, and experimental results demonstrate that enabling AGQ can achieve an up to 100% increase in training throughput without explicit drop of training accuracy compared with existing INA solutions like ATP and host-based quantization methods like THC.
Huifeng Xing, Yinfan Hu, Hao Wang 0231, Yang Chen 0001, Sen Liu 0002, Yang Xu 0010
ICDCS2
2024 Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model Training
abstract
Parameter Server (PS) and Ring-AllReduce (RAR) are two widely utilized synchronization architectures in multiworker Deep Learning (DL), also referred to as Distributed Deep Learning (DDL). However, PS encounters challenges with the “incast” issue, while RAR struggles with problems caused by the long dependency chain. The emerging In-network Aggregation (INA) has been proposed to integrate with PS to mitigate its incast issue. However, such PS-based INA has poor incremental deployment abilities as it requires replacing all the switches to show significant performance improvement, which is not costeffective. In this study, we present the incorporation of INA capabilities into RAR, called RAR with In-Network Aggregation (Rina), to tackle both the problems above. Rina features its agent-worker mechanism. When an INA-capable ToR switch is deployed, all workers in this rack run as one abstracted worker with the help of the agent, resulting in both excellent incremental deployment capabilities and better throughput. We conducted extensive testbed and simulation evaluations to substantiate the throughput advantages of Rina over existing DDL training synchronization structures. Compared with the state-of-the-art PS-based INA methods ATP, Rina can achieve more than$\mathbf{5 0 \%}$throughput with the same hardware cost.
Xuandong Liu, Minglin Li, Yinfan Hu, Huifeng Xing, Hao Wang 0231, Wanxin Shi, Sen Liu 0002, Yang Xu 0010
ICNP4