VLDB 2026 Research / reviewers in the wild / expert
Yang Zhang 0103
dblp:06/6785-103
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0002-3965-2949ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call AutomationabstractIn large-scale enterprises, on-call engineers (OCEs) are critical for ensuring service availability and reliability. However, as incidents grow in volume and complexity, traditional manual on-call processes are becoming increasingly inadequate. Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in reasoning and multi-agent collaboration, presenting new opportunities for automation. We propose OncallX, an end-to-end automated on-call system designed for real-world industrial scenarios that integrates LLMs with multi-agent cooperation to enable intelligent and efficient incident management. OncallX first enhances user queries by leveraging external knowledge bases and multi-turn dialogue interactions. Subsequently, multiple expert agents collaborate through tree-search-based mechanisms to generate effective responses and solutions. When incidents cannot be resolved automatically, OncallX accurately assigns them to the most appropriate teams. Comprehensive experiments conducted in the real-world production environment of a top-tier global online video service provider demonstrate that OncallX efficiently responds to incidents and accurately triages tickets, significantly outperforming existing methods in both automated metrics and human evaluations. Furthermore, OncallX has been successfully deployed in production for two months, during which it has substantially enhanced on-call efficiency, reducing average incident response time to just 21 seconds and average triage time to 4 seconds—representing a transformative improvement in operational excellence. Ruowei Fu, Yang Zhang 0103, Zeyu Che, Zhenyu Zhong, Zhiqiang Ren, Shenglin Zhang, Feng Wang 0054, Yongqian Sun, Yu Zhang 0209 |
ASE | 2 |
| 2024 | G-Cause: Parameter-free Global Diagnosis for Hyperscale Web Service InfrastructuresabstractHyperscale web service infrastructures are becoming increasingly complex and facing a variety of threats, raising the demand for more sophisticated automated operations and diagnosis solutions. Existing anomaly root cause localization approaches often focus on Service-level components without drilling down to the lower-level resources where services are deployed, hindering the implementation of fine-grained failure fix measures. This paper introduces a challenging task called global diagnosis and addresses it by proposing a technique called G-Cause, which is applicable to both Service-level and host-level root cause analysis scenarios. G-Cause builds a highly adaptive diagnostic framework based on the frequency-domain and time-domain characteristics of monitoring metrics, allowing it to handle global diagnosis requirements from app to host with minimal parameter adjustments. We deploy and validate our approach in two typical scenarios: homogeneous metric diagnosis from app to microservice, and heterogeneous metric diagnosis for various host resources. The results demonstrate that G-Cause outperforms state-of-the-art diagnosis algorithms while providing strong interpretability. Our approach helps operators understand the core mechanism of anomaly propagation and adjust their management strategies more effectively. With these strengths, G-Cause successfully services our global product operations and also makes an impressive contribution in many other workflows. Xinrui Jiang 0001, Yang Zhang 0103, Tingzhu Bi, Xiangzhuang Shen, Yu Zhang 0209, Yicheng Pan 0002, Meng Ma 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003 |
ICWS | 2 |
| 2024 | FaultInsight: Interpreting Hyperscale Data Center Host FaultsabstractOperating and maintaining hyperscale data centers involving millions of service hosts has been an extremely intricate task to tackle for top Internet companies.Incessant system failures cost operators countless hours of browsing through performance metrics to diagnose the underlying root cause to prevent the recurrence.Although many state-of-the-art (SOTA) methods have used time-series causal discovery to construct causal relationships among anomalous metrics, they only focus on homogeneous service-level performance metrics and fail to yield useful insights on heterogeneous host-level metrics.To address the challenge, this study presents FaultInsight, a highly interpretable deep causal host fault diagnosing framework that offers diagnostic insights from various perspectives to reduce human effort in troubleshooting.We evaluate FaultInsight using dozens of incidents collected from our production environment.FaultInsight provides markedly better root cause identification accuracy than SOTA baselines in our incident dataset.It also shows outstanding advantages in terms of deployability in real production systems.Our engineers are deeply impressed by FaultInsight's ability to interpret incidents from multiple perspectives, helping them quickly understand the mechanism behind the faults. Tingzhu Bi, Yang Zhang 0103, Yicheng Pan 0002, Yu Zhang 0209, Meng Ma 0001, Xinrui Jiang 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003 |
KDD | 2 |