Yongqi Han 0001

dblp:300/8499 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0003-1244-0663ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Semi-Supervised Metrics-Based Self-Training Root Cause Analysis for Cloud-Native Systems with Class-Imbalanced Data
abstract
Root cause analysis is crucial for cloud-native systems. However, existing supervised approaches ignore the potential of unlabeled data, which is frequent in the cloud-native root cause analysis scenarios. Moreover, the class-imbalanced distribution of faults presents obstacles to applying semi-supervised learning. To overcome these limitations, we propose STRCA, a metrics-based semi-supervised self-training approach for root cause analysis. Furthermore, STRCA employs minority priority self-training, which selects pseudo-labels of high quality during generations. Additionally, the stepwise distribution alignment is introduced to rebalance the predicted distribution with gradually decreasing strength. These two strategies mitigate the class-imbalance of data in semi-supervised learning. Experiments on the public dataset show the effectiveness of STRCA with limited labels.
Qingfeng Du, Yongqi Han 0001, Fulong Tian
ICASSP3
2024 The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small Classifier
abstract
Failure root cause analysis (RCA), which systematically identifies underlying faults, is essential for ensuring the reliability of widely adopted microservice-based applications and cloud-native systems. However, manual analysis by simple rules faces significant burdens due to the heterogeneous nature of resource entities and the massive amount of observability data. Furthermore, existing approaches for automating RCA struggle to perform in-depth fault analysis without extensive fault labels. To address the scarcity of fault labels, we examine an extreme RCA scenario where each fault type has only one example (one-shot). We propose LasRCA, a framework for one-hot RCA in cloud-native systems that leverages the collaboration of the large language model (LLM) and the small classifier. In the training stage, LasRCA initially trains a small classifier based on one-shot fault examples. The small classifier then iteratively selects high-confusion samples and receives feedback on their fault types from LLM-driven fault labeling. These samples are applied to retrain the small classifier. In the inference stage, LasRCA performs a joint RCA through the collaboration of the LLM and small classifier, achieving a trade-off between effectiveness and cost. Experiment results on public datasets with heterogeneous nature and prevalent fault types show the effectiveness of LasRCA in one-shot RCA.
Yongqi Han 0001, Qingfeng Du, Fulong Tian
ASE1
2024 Holistic Root Cause Analysis for Failures in Cloud-Native Systems Through Observability Data
abstract
Microservices are widely adopted in large IT enterprises, leveraging the scalability, resiliency, and elasticity of the cloud-native architecture. Effective root cause analysis is crucial for ensuring the reliability of such cloud-native systems. Many efforts have focused on using the three modalities of observability data–traces, metrics, and logs. However, existing approaches are limited by inconsistent problem definitions and cloud-native heterogeneity. To address these challenges, we proposeHolisticRCA, a root cause analysis framework in cloud-native systems from a holistic perspective.HolisticRCAformally defines root cause analysis through three dimensions. ThenHolisticRCAuses an “assembling building blocks” strategy to address the cloud-native heterogeneity. It maps each observability feature into a shared vector space and concatenates the vector embeddings associated with each resource entity for standardized resource entity vector embeddings. Then it applies Graph Attention Network to capture intertwined resource entity relations and incorporates mask embeddings to enable holistic analysis. The evaluation results on three public datasets show thatHolisticRCAoutperforms existing approaches in holistic root cause analysis of cloud-native systems.
Yongqi Han 0001, Qingfeng Du, Pengsheng Li, Xiaonan Shi, Pei Fang, Fulong Tian
IEEE Trans. Serv. Comput.1
2023 Trace-Based Anomaly Detection with Contextual Sequential Invocations
Qingfeng Du, Fulong Tian, Yongqi Han 0001
DEXA (2)4
2023 LWS: A framework for log-based workload simulation in session-based SUT
Yongqi Han 0001, Qingfeng Du, Jincheng Xu, Shengjie Zhao 0001, Zhekang Chen, Kanglin Yin, Dan Pei
J. Syst. Softw.1
2021 Log-Based Anomaly Detection with Multi-Head Scaled Dot-Product Attention Mechanism
Qingfeng Du, Jincheng Xu, Yongqi Han 0001, Shuangli Zhang
DEXA (1)4