Yuhong Liao

dblp:417/1518 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0003-7216-5592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Network management and operations · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network management and operations › fault management
fault diagnosis
1.722025
SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures · SIGCOMM 2025
Towards LLM-Based Failure Localization in Production-Scale Networks · SIGCOMM 2025
Network management and operations › fault management
failure localization
0.912025
Towards LLM-Based Failure Localization in Production-Scale Networks · SIGCOMM 2025
Network management and operations
network restoration
0.912025
SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures · SIGCOMM 2025
Network management and operations › fault management › fault diagnosis
root cause analysis
0.912025
Towards LLM-Based Failure Localization in Production-Scale Networks · SIGCOMM 2025

Methods — techniques the papers use, named apart from their topics

severity assessment · 1.7alert grouping · 1.7large language model · 0.9
YearPublicationVenuePosition
2025 Towards LLM-Based Failure Localization in Production-Scale Networks
abstract
Root causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it extremely challenging even for experienced operators. Large language models (LLMs) have shown great potential in text understanding and reasoning. In this paper, we present BiAn, an LLM-based framework designed to assist operators in efficient incident investigation. BiAn processes monitoring data and generates error device rankings with detailed explanations. To date, BiAn has been deployed in our network infrastructure for 10 months and it has successfully assisted operators in identifying error devices more quickly, reducing time to root causing by 20.5% (55.2% for high-risk incidents). Extensive performance evaluations based on 17 months of real cases further demonstrate that BiAn achieves accurate and fast failure localization. It improves accuracy by 9.2% compared to the baseline approach.
Chenxu Wang 0007, Xumiao Zhang, Runwei Lu, Xianshang Lin, Xuan Zeng 0002, Zhe An, Gongwei Wu, Chen Tian 0001, Guihai Chen, Guyue Liu, Yuhong Liao, Dennis Cai, Ennan Zhai
SIGCOMM13
2025 SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures
abstract
For providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network failures. In practice, there is a gap between the flooding raw alerts data collected by network monitoring tools and the readable information needed for failure diagnosis. Existing solutions using limited network monitoring data sources and heuristic diagnostic rules, lack comprehensive coverage and the capability to address severe failures, especially which network operators have never handled a similar one before. This paper presents SkyNet, a network analysis system to extract scope and severity information from alert floods. SkyNet ensures comprehensive coverage by integrating multiple monitoring data sources through a uniform input format, enhancing extensibility for new network monitoring tools. During alert floods, SkyNet groups alerts, assesses their severity, and filters out insignificant ones to aid network operators in mitigating network failures. To date, SkyNet has been running stably on our network for one and a half years without any false negatives and has successfully reduced the time-to-mitigation for over 80% of network failures since its deployment in production.
Huanwu Hu, Yunguang Li, Xiangyu Tang, Bingchuan Tian, Gongwei Wu, Xumiao Zhang, Ennan Zhai, Yuhong Liao, Dennis Cai
SIGCOMM13