Shixian Guo

dblp:379/4191 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-0542-8281ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Skyline: A Cloud Centric Internet Monitoring Engine
Shixian Guo, Yangyang Bai, Kefei Liu 0004, Zhenyang Zhong, Sisi Wen, Yongbin Dong, Anjian Chen, Jiale Feng, Lingpei Meng, Siwan Chen, Juntao Zhong, Chaoran Hu, Yibo Huang 0005, Yiming Qiu 0001
NSDI1
2026 Horizon: A Hyper-Edge Observability Engine for Live Streaming Networks
abstract
Live streaming services power mainstream real-time interactions on top of dedicated live streaming networks (LiveNets). Yet making LiveNets reliable at scale is challenging: failures arise on the userfacing delivery path and within streaming protocol and application logic, so operators need both continuous runtime monitoring to detect and localize incidents quickly and proactive preflight testing to exercise changes under representative environments and sustained playback behavior. Meeting these goals hinges on the right vantage point: the observability workflow must traverse the same network paths and delivery stacks as users while remaining controllable and non-intrusive. We present Horizon, which leverages near-user, provider-managed hyper-edge devices and orchestrates them into a shared fleet that supports both always-on monitoring and customizable, scenario-driven validation. Horizon has been deployed in production for over three years; in 2025, it identified 2,000+ major network incidents using 100,000+ hyper-edge agents.
Daqian Ding, Shixian Guo, Zhendong Xie, Aifang Xu, Changqian Wang, Kefei Liu 0004, Jialin Li 0001, Yunming Xiao, Heming Cui, Yiming Qiu 0001
SIGCOMM4
2025 ByteTracker: An Agentless and Real-time Path-aware Network Probing System
abstract
As the number of data center servers grows into the millions and due to the demand for more accurate, rapid and powerful network fault detection and location, the existing Pingmesh-centric monitoring and diagnostic system is not efficient enough. In this paper, we propose ByteTracker, the first agentless probing and diagnostic system for large-scale data center networks. It does not need to deploy probe processes or make any configurations on end hosts, and all probes are launched by a small number of centralized Probers. ByteTracker achieves accurate, real-time probe path tracking with packet mirroring on switches. By reducing end-host probe noise, precisely identifying network timeout probes, accurately tracking probe paths, and marking the failed switch with multiple network timeout probes, ByteTracker can locate network failures with nearly 100% accuracy. We have deployed ByteTracker in all of our data centers for over half a year. During deployment, ByteTracker can detect almost all network anomalies and locate them within 5 seconds with 100% accuracy.
Shixian Guo, Kefei Liu 0004, Yulin Lai, Yangyang Bai, Jianghang Ning, Yongbin Dong, Sisi Wen, Jiale Feng, Chengcai Yao, Zhuo Jiang, Jiao Zhang 0002, Tao Huang 0005
SIGCOMM1
2024 Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE Clusters
abstract
RoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary.
Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Zicheng Wang 0004, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
APNet5
2024 R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic System
abstract
RoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related.
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Haohan Xu, Dongyang Song, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
SIGCOMM4