Sibo Xia

dblp:312/5032 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A time-frequency dual-branch feature dynamic fusion prediction network for tail gas sulfur content prediction in the wet flue gas desulfurization process
Siheng Zeng, Hongqiu Zhu, Sibo Xia, Bochun Yue
Eng. Appl. Artif. Intell.3
2026 HHMODE: a global-local co‑evolutionary hyper‑heuristic algorithm for coordinated pump scheduling in urban water distribution systems
abstract
In urban water distribution systems, the scientific design of pump group scheduling strategies directly affects both water supply safety and operational cost-effectiveness. However, the operation of pump groups involves coupling conflicts among multiple objectives and complex operational constraints, making it difficult for traditional methods to effectively solve the scheduling optimization problem. Meanwhile, the continuous rise in urban water demand and energy prices introduces new challenges to achieving efficient and energy-saving pump operation. To address this issue, a multi-objective collaborative pump group scheduling model is proposed, fully considering the regulation function of the clear water pool. Based on this, a global–local co-evolutionary hyper-heuristic algorithm is developed. In the intake-supply collaborative optimization model, an adaptive time-division strategy is designed for intake flow planning, which reduces the load on intake pumps by smoothing flow distribution. In the hyper-heuristic optimization algorithm, a hybrid adaptive triggering mechanism is designed to dynamically coordinate the cooperative co-evolution of global optimization and local search. Meanwhile, a multi-armed bandit strategy is employed to adaptively select local search operators online, thereby efficiently identifying the best individuals within promising search regions. Validated on a real-world pump scheduling case, the proposed algorithm achieves superior optimization performance, reducing total energy consumption by 20.02% and decreasing the number of pump switching operations by 44% compared with the on-site scheduling strategy, and it is capable of producing accurate pump scheduling plans within a limited number of iterations. Furthermore, it exhibits superior performance on standard benchmark functions, demonstrating strong applicability.
Sibo Xia, Hongqiu Zhu, Hangmin Zhao
Expert Syst. Appl.1
2026 Failure Diagnosis in Microservice Systems: A Comprehensive Survey and Analysis
abstract
Widely adopted for their scalability and flexibility, modern microservice systems present unique failure diagnosis challenges due to their independent deployment and dynamic interactions. This complexity can lead to cascading failures that negatively impact operational efficiency and user experience. Recognizing the critical role of fault diagnosis in improving the stability and reliability of microservice systems, researchers have conducted extensive studies and achieved a number of significant results. This survey provides an exhaustive review of 98 scientific papers from 2003 to the present, including a thorough examination and elucidation of the fundamental concepts, system architecture, and problem statement. It also includes a qualitative analysis of the dimensions, providing an in-depth discussion of current best practices and future directions, aiming to further its development and application. In addition, this survey compiles publicly available datasets, toolkits, and evaluation metrics to facilitate the selection and validation of techniques for practitioners.
Shenglin Zhang, Sibo Xia, Wenzhao Fan, Binpeng Shi, Xiao Xiong, Zhenyu Zhong, Minghua Ma, Yongqian Sun, Dan Pei
ACM Trans. Softw. Eng. Methodol.2
2025 Efficient and Accurate Anomaly Detection in HPC Systems via Coarse-Grained Clustering and Fine-Grained Model Sharing
Shaoyu Hu, Sibo Xia, Yongqian Sun, Xijie Pan, Yuan Yuan 0034, Shenglin Zhang
APNet2
2025 Forewarned is Forearmed: Joint Prediction and Classification of Optical Transceiver Failures in Large-Scale LLM Training Clusters
Sibo Xia, Junhua Kuang, Shenglin Zhang, Qitong Xie, Yongqian Sun
APNet1
2025 Effective Node-Level Anomaly Detection in HPC Systems via Coarse-Grained Clustering and Fine-Grained Model Sharing
abstract
High-performance computing (HPC) systems are crucial for scientific advancement and engineering breakthroughs. Unexpected performance degradation or system failures can severely impact these endeavors. This paper introduces NodeSentry, a novel unsupervised anomaly detection framework tailored for compute nodes of large-scale HPC systems. NodeSentry leverages a combined approach of coarse-grained clustering and fine-grained model sharing to effectively address the challenges posed by the massive node scales, frequent job transitions, and complex patterns characteristic of modern HPC deployments. Evaluation on two real-world HPC datasets demonstrates NodeSentry’s superior performance, achieving an F1-score exceeding 0.876. This represents a 0.560 average improvement over existing best baseline methods, while simultaneously reducing training overhead by an average of 45.69%. Furthermore, to promote reproducibility and contribute to the broader research community, we open-source NodeSentry’s codebase and introduce a novel clustering adjustment and anomaly labeling tool specifically designed for HPC systems.
Sibo Xia, Yongqian Sun, Xijie Pan, Yuan Yuan 0034, Shenglin Zhang, Shaoyu Hu, Jinghua Feng
SC1
2024 ART: A Unified Unsupervised Framework for Incident Management in Microservice Systems
abstract
Automated incident management is critical for large-scale microservice systems, including tasks such as anomaly detection (AD), failure triage (FT), and root cause localization (RCL). Currently, most techniques focus only on a single task, overlooking shared knowledge across closely related tasks. However, employing isolated models for managing multiple tasks may result in inefficiencies, delayed responses, a lack of systemic perspective, and complexity in updates and operations. Therefore we propose ART, an unsupervised framework that integrates a full-process solution covering Anomaly detection, failure Triage, and Root cause localization. It reaches the unification of multiple tasks by extracting the shared knowledge. Specifically, we first conduct an empirical study to analyze how the shared knowledge embedded in anomalous deviations manifests in AD, FT, and RCL. To better calculate deviations and extract shared knowledge, we sequentially model channel, temporal, and call dependencies using Transformer Encoder, GRU, and GraphSAGE, respectively. Then unified failure representations enhance the interpretability of abstract features with explicit semantic information, serving as the basis for unsupervised multitask solutions. Our evaluations on the datasets generated from two benchmark microservice systems demonstrate that ART outperforms existing methods in terms of AD (improving by 5.65% to 60.8%), FT (improving by 13.2% to 95.7%), and RCL (improving by 13.3% to 205%).
Yongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma, Sibo Xia, Shenglin Zhang, Dan Pei
ASE5
2024 DBFiLM: A novel dual-branch frequency improved legendre memory forecasting model for coagulant dosage determination
Sibo Xia, Hongqiu Zhu, Yonggang Li 0002, Can Zhou 0005
Expert Syst. Appl.1
2024 No More Data Silos: Unified Microservice Failure Diagnosis With Temporal Knowledge Graph
abstract
Microservices improve the scalability and flexibility of monolithic architectures to accommodate the evolution of software systems, but the complexity and dynamics of microservices challenge system reliability. Ensuring microservice quality requires efficient failure diagnosis, including detection and triage. Failure detection involves identifying anomalous behavior within the system, while triage entails classifying the failure type and directing it to the engineering team for resolution. Unfortunately, current approaches reliant on single-modal monitoring data, such as metrics, logs, or traces, cannot capture all failures and neglect interconnections among multimodal data, leading to erroneous diagnoses. Recent multimodal data fusion studies struggle to achieve deep integration, limiting diagnostic accuracy due to insufficiently captured interdependencies. Therefore, we proposeUniDiag, which leverages temporal knowledge graphs to fuse multimodal data for effective failure diagnosis.UniDiagapplies a simple yet effective stream-based anomaly detection method to reduce computational cost and a novel microservice-oriented graph embedding method to represent the state of systems comprehensively. To assess the performance ofUniDiag, we conduct extensive evaluation experiments using datasets from two benchmark microservice systems, demonstrating its superiority over existing methods and affirming the efficacy of multimodal data fusion. Additionally, we have publicly made the code and data available to facilitate further research.
Shenglin Zhang, Sibo Xia, Shirui Wei, Yongqian Sun, Shiyu Ma, Junhua Kuang, Bolin Zhu, Lemeng Pan, Yicheng Guo, Dan Pei
IEEE Trans. Serv. Comput.3
2023 Robust Failure Diagnosis of Microservice System Through Multimodal Data
abstract
Automatic failure diagnosis is crucial for large microservice systems. Currently, most failure diagnosis methods rely solely on single-modal data (i.e., using either metrics, logs, or traces). In this study, we conduct an empirical study using real-world failure cases to show that combining these sources of data (multimodal data) leads to a more accurate diagnosis. However, effectively representing these data and addressing imbalanced failures remain challenging. To tackle these issues, we proposeDiagFusion, a robust failure diagnosis approach that uses multimodal data. It leverages embedding techniques and data augmentation to represent the multimodal data of service instances, combines deployment data and traces to build a dependency graph, and uses a graph neural network to localize the root cause instance and determine the failure type. Our evaluations using real-world datasets show thatDiagFusionoutperforms existing methods in terms of root cause instance localization (improving by 20.9% to 368%) and failure type determination (improving by 11.0% to 169%).
Shenglin Zhang, Pengxiang Jin, Yongqian Sun, Bicheng Zhang, Sibo Xia, Zhengdan Li, Zhenyu Zhong, Minghua Ma, Wa Jin, Dan Pei
IEEE Trans. Serv. Comput.6
2022 Robust System Instance Clustering for Large-Scale Web Services
abstract
System instance clustering is crucial for large-scale Web services because it can significantly reduce the training overhead of anomaly detection methods. However, the vast number of system instances with massive time points, redundant metrics, and noise bring significant challenges. We propose OmniCluster to accurately and efficiently cluster system instances for large-scale Web services. It combines a one-dimensional convolutional autoencoder (1D-CAE), which extracts the main features of system instances, with a simple, novel, yet effective three-step feature selection strategy. We evaluated OmniCluster using real-world data collected from a top-tier content service provider providing services for one billion+ monthly active users (MAU), proving that OmniCluster achieves high accuracy (NMI=0.9160) and reduces the training overhead of five anomaly detection models by 95.01% on average.
Shenglin Zhang, Dongwen Li, Zhenyu Zhong, Minghan Liang, Jiexi Luo, Yongqian Sun, Ya Su, Sibo Xia, Zhongyou Hu, Dan Pei, Jiyan Sun, Yinlong Liu
WWW9