Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yibo Xiao

dblp:362/9861 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0002-9036-4563ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
3 papers
Software-defined and programmable networks · 36% Internet architecture and protocols · 23% Datacenter networks · 16%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 62% Performance modeling and evaluation · 23% Interconnection networks and networks-on-chip · 15%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
anomaly detection
1.012026
Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement · ASPLOS (2) 2026
Distributed systems
fault tolerance
1.012026
Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement · ASPLOS (2) 2026
Performance modeling and evaluation
performance monitoring
1.012026
Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement · ASPLOS (2) 2026
Datacenter networks
RDMA
1.022026
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023
Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement · ASPLOS (2) 2026
Internet of things and sensor networks › wireless sensor network
network diagnosis
0.912025
Troubleshooting Programmable Data Planes via Real-Time Table Information Recording · IEEE Trans. Netw. 2025
Software-defined and programmable networks
programmable data plane
0.912025
Troubleshooting Programmable Data Planes via Real-Time Table Information Recording · IEEE Trans. Netw. 2025
Internet architecture and protocols › multicast
in-network multicast
0.712023
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023
Internet architecture and protocols › multicast
reliable multicast
0.712023
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023
Distributed systems › replication
data replication
0.712023
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023
Interconnection networks and networks-on-chip › remote direct memory access
RDMA-based replication
0.712023
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023
Natural language and speech › Language models and text generation
large language model training
0.312026
Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement · ASPLOS (2) 2026
Network measurement and analytics
traffic measurement
0.312026
Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement · ASPLOS (2) 2026
Network management and operations › fault management
fault diagnosis
0.312025
Troubleshooting Programmable Data Planes via Real-Time Table Information Recording · IEEE Trans. Netw. 2025
Software-defined and programmable networks › programmable data plane
p4 switch
0.212023
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023
Software-defined and programmable networks › programmable data plane
programmable switch
0.212023
MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support · ICNP 2023

Methods — techniques the papers use, named apart from their topics

microsecond-level traffic measurement · 3.0RDMA flow measurement · 3.0multicast routing protocol · 1.3ACK/NAK merging · 1.3table recording · 0.9probabilistic transition DAG · 0.9information entropy · 0.9
YearPublicationVenuePosition
2026 Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic Measurement
abstract
Large language model (LLM) training is prone to anomalies due to its long duration and large scale, which can lead to significant performance degradation or even training crashes. Due to the synchronization nature of LLM training, anomalies exhibit the cascading effect, making their diagnosis challenging. Existing approaches rely on collecting communication operator information via code instrumentation, which yields only coarse-grained monitoring data and requires modifications to training code or communication libraries. We propose Pulse, a fine-grained, non-intrusive, and easy-to-deploy monitoring system. Our key idea is to enable fine-grained monitoring via traffic measurement. Pulse conducts microsecond-level RDMA traffic measurement on NICs, and transforms flow-level measurements into communication operator measurements, thereby enabling fine-grained and non-intrusive monitoring. We deploy Pulse on a testbed with 64 H200 GPUs and evaluate its anomaly localization capability under common failure scenarios. Pulse achieves machine-level localization in 10 out of 12 scenarios, while existing methods succeed in only 4 and even misdiagnose 2 of the remaining scenarios. Additionally, Pulse achieves over 90% precision and 100% recall, supports up to 2000 concurrent RDMA flow measurements per NIC, and imposes negligible overhead on training performance, making it a practical solution for real-world LLM training environments.
Yibo Xiao, Haifeng Sun 0004, Qingkai Meng 0001, Jiong Duan, Xiaohe Hu, Rong Gu 0001, Guihai Chen, Chen Tian 0001
ASPLOS (2)1
2025 Troubleshooting Programmable Data Planes via Real-Time Table Information Recording
abstract
While the flexibility of programmable switches brings opportunities, it also introduces security risks. Hence, it is vital to conduct effective troubleshooting in the programmable switch to mitigate frequent network failures. However, troubleshooting programmable switch failures is challenging due to their enhanced flexibility and functionality compared to regular switches, posing increased difficulty in debugging, particularly with limited debugging tools and information. To address this problem, we propose an efficient troubleshooting method that records real-time information about packets in the data plane, including the tables involved in packet processing. Unfortunately, due to hardware limitations, it is infeasible to record all tables’ information in the data plane. Thus, the key is to find the table set reflecting the execution path a packet goes through while minimizing the resource overhead. We first represent P4 programs as a probabilistic transition directed acyclic graph (DAG) and employ information entropy to quantify the information within a set of tracked tables. Then, we adopt a two-step approach and design algorithms to find both optimal and approximately optimal table record plans. The evaluation results show the efficacy of the proposed method, including achieving the same path recovery rate as the related works with less than one-third of the resource consumption.
Chengyuan Huang, Yibo Xiao, Tianfan Zhang, Bingheng Yan, Ahmed M. Abdelmoniem, Gianni Antichi, Xiaoliang Wang 0001, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.2
2024 MpScope: Enabling multi-pipeline monitoring inside a switch
Chengyuan Huang, Tianfan Zhang, Li Wang 0110, Yibo Xiao, Chen Tian 0001, Xiaoliang Wang 0001, Bingheng Yan, Ahmed M. Abdelmoniem, Wan-Chun Dou, Guihai Chen
Comput. Networks4
2023 MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support
abstract
Remote Direct Memory Access has been widely adopted in distributed storage systems. However, it only supports unicast operations, which degrades the performance significantly for data replication because of bandwidth waste and CPU overhead. To address the problem, we propose MC-RDMA, a distributed and reliable multicast RDMA. It is compatible with existing unicast RDMA but supports lazy packet replication with reliable RDMA multicasting. The key idea of MC-RDMA is utilizing in-network programmable switches to build a NIC-transparent reliable multicast protocol for RDMA. MC-RDMA combines the address information of the IP and RoCEv2 into a sender-initialized multicast routing protocol. Besides, it synchronizes the hardware transmission states of multiple receivers by merging ACKs and NAKs. To verify the effectiveness of MC-RDMA, we implement it with Mellanox ConnectX-6 commodity RNICs and Intel Tofino P4 programmable switches. Experimental results show that MC-RDMA can double the sender bandwidth utilization and reduce the CPU overhead significantly compared to unicast-based RDMA replications. Moreover, it reduces the storage request latency by -30% with realistic workloads and decreases the training time by -50% in the distributed training system.
Chengyuan Huang, Yixiao Gao, Duoxing Li, Yibo Xiao, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Fu Xiao 0001
ICNP5