Yongbin Dong

dblp:373/3740 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Skyline: A Cloud Centric Internet Monitoring Engine
Shixian Guo, Yangyang Bai, Kefei Liu 0004, Zhenyang Zhong, Sisi Wen, Yongbin Dong, Anjian Chen, Jiale Feng, Lingpei Meng, Siwan Chen, Juntao Zhong, Chaoran Hu, Yibo Huang 0005, Yiming Qiu 0001
NSDI13
2025 ByteTracker: An Agentless and Real-time Path-aware Network Probing System
abstract
As the number of data center servers grows into the millions and due to the demand for more accurate, rapid and powerful network fault detection and location, the existing Pingmesh-centric monitoring and diagnostic system is not efficient enough. In this paper, we propose ByteTracker, the first agentless probing and diagnostic system for large-scale data center networks. It does not need to deploy probe processes or make any configurations on end hosts, and all probes are launched by a small number of centralized Probers. ByteTracker achieves accurate, real-time probe path tracking with packet mirroring on switches. By reducing end-host probe noise, precisely identifying network timeout probes, accurately tracking probe paths, and marking the failed switch with multiple network timeout probes, ByteTracker can locate network failures with nearly 100% accuracy. We have deployed ByteTracker in all of our data centers for over half a year. During deployment, ByteTracker can detect almost all network anomalies and locate them within 5 seconds with 100% accuracy.
Shixian Guo, Kefei Liu 0004, Yulin Lai, Yangyang Bai, Jianghang Ning, Yongbin Dong, Sisi Wen, Jiale Feng, Chengcai Yao, Zhuo Jiang, Jiao Zhang 0002, Tao Huang 0005
SIGCOMM10
2024 Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE Clusters
abstract
RoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary.
Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Zicheng Wang 0004, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
APNet7
2024 NetAssistant: Dialogue Based Network Diagnosis in Data Center Networks
Haopei Wang, Anubhavnidhi Abhashkumar, Changyu Lin, Tianrong Zhang, Xiaoming Gu, Yongbin Dong, Weirong Jiang
NSDI10
2024 R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic System
abstract
RoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related.
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Haohan Xu, Dongyang Song, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
SIGCOMM7