Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yongqing Xi

dblp:271/5959 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0007-0520-2196ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Datacenter networks · 48% Network measurement and analytics · 27% Software-defined and programmable networks · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 77% Storage systems · 23%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Datacenter networks
data center network topology
0.812024
Alibaba HPN: A Data Center Network for Large Language Model Training · SIGCOMM 2024
Datacenter networks
load balancing
0.812024
Alibaba HPN: A Data Center Network for Large Language Model Training · SIGCOMM 2024
Cloud and datacenter computing
cloud storage
0.512021
When Cloud Storage Meets RDMA · NSDI 2021
Network measurement and analytics › network telemetry
in-network telemetry
0.412020
Flow Event Telemetry on Programmable Data Plane · SIGCOMM 2020
Software-defined and programmable networks
programmable data plane
0.412020
Flow Event Telemetry on Programmable Data Plane · SIGCOMM 2020
Network management and operations › network robustness
fault tolerance
0.212024
Alibaba HPN: A Data Center Network for Large Language Model Training · SIGCOMM 2024
Storage systems › networked storage › storage networking
RDMA storage
0.112021
When Cloud Storage Meets RDMA · NSDI 2021

Methods — techniques the papers use, named apart from their topics

message batching · 0.4information compression · 0.4event detection · 0.4
YearPublicationVenuePosition
2026 Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001
SIGCOMM19
2024 Alibaba HPN: A Data Center Network for Large Language Model Training
abstract
This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production.
Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai
SIGCOMM2
2021 When Cloud Storage Meets RDMA
Yixiao Gao, Qiang Li 0045, Lingbo Tang, Yongqing Xi, Wenwen Peng, Bo Li 0061, Yaohui Wu, Shaozong Liu, Xingkui Liu, Zhongjie Wu, Junping Wu, Zheng Cao 0003, Chen Tian 0001, Jiaji Zhu, Haiyong Wang, Dennis Cai, Jiesheng Wu
NSDI4
2020 Flow Event Telemetry on Programmable Data Plane
abstract
Network performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic.
Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001
SIGCOMM10