Yicheng Pan 0002

dblp:14/721-2 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-4139-1477ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PowerCause: Leveraging Causality for Self-Recovery in Unbalanced Distribution Networks
abstract
This article proposes PowerCause, a method for phase unbalance positioning and active regulation in smart distribution networks. PowerCause can automatically detect anomaly intervals, locate the source of phase unbalance using Granger causality test, back-search, and generate a list of potential root cause buses. It takes regulation measures for specific buses to alleviate the impact of the unbalance. This implements a closed-loop solution that handles the entire process from the occurrence of phase unbalance to its positioning, and finally, to active regulation. This study builds a closed-loop simulation environment based on the open distribution system simulator (OpenDSS) to enable autonomous and controllable unbalance injection, collect multidimensional bus metrics, including voltage and phase angle. The environment also includes causal analysis and active regulation modules for method verification. The proposed method demonstrates high accuracy in root cause location, efficient performance, and robustness against environmental influences such as measurement noise and data loss errors.
Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003
IEEE Trans. Reliab.2
2025 UnCLe: Towards Scalable Dynamic Causal Discovery in Non-linear Temporal Systems
abstract
Uncovering cause-effect relationships from observational time series is fundamental to understanding complex systems. While many methods infer static causal graphs, real-world systems often exhibit *dynamic causality*—where relationships evolve over time. Accurately capturing these temporal dynamics requires time-resolved causal graphs. We propose UnCLe, a novel deep learning method for scalable dynamic causal discovery. UnCLe employs a pair of Uncoupler and Recoupler networks to disentangle input time series into semantic representations and learns inter-variable dependencies via auto-regressive Dependency Matrices. It estimates dynamic causal influences by analyzing datapoint-wise prediction errors induced by temporal perturbations. Extensive experiments demonstrate that UnCLe not only outperforms state-of-the-art baselines on static causal discovery benchmarks but, more importantly, exhibits a unique capability to accurately capture and represent evolving temporal causality in both synthetic and real-world dynamic systems (e.g., human motion). UnCLe offers a promising approach for revealing the underlying, time-varying mechanisms of complex phenomena.
Tingzhu Bi, Yicheng Pan 0002, Xinrui Jiang 0001, Huize Sun, Meng Ma 0001, Ping Wang 0003
NeurIPS2
2024 G-Cause: Parameter-free Global Diagnosis for Hyperscale Web Service Infrastructures
abstract
Hyperscale web service infrastructures are becoming increasingly complex and facing a variety of threats, raising the demand for more sophisticated automated operations and diagnosis solutions. Existing anomaly root cause localization approaches often focus on Service-level components without drilling down to the lower-level resources where services are deployed, hindering the implementation of fine-grained failure fix measures. This paper introduces a challenging task called global diagnosis and addresses it by proposing a technique called G-Cause, which is applicable to both Service-level and host-level root cause analysis scenarios. G-Cause builds a highly adaptive diagnostic framework based on the frequency-domain and time-domain characteristics of monitoring metrics, allowing it to handle global diagnosis requirements from app to host with minimal parameter adjustments. We deploy and validate our approach in two typical scenarios: homogeneous metric diagnosis from app to microservice, and heterogeneous metric diagnosis for various host resources. The results demonstrate that G-Cause outperforms state-of-the-art diagnosis algorithms while providing strong interpretability. Our approach helps operators understand the core mechanism of anomaly propagation and adjust their management strategies more effectively. With these strengths, G-Cause successfully services our global product operations and also makes an impressive contribution in many other workflows.
Xinrui Jiang 0001, Yang Zhang 0103, Tingzhu Bi, Xiangzhuang Shen, Yu Zhang 0209, Yicheng Pan 0002, Meng Ma 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003
ICWS6
2024 FaultInsight: Interpreting Hyperscale Data Center Host Faults
abstract
Operating and maintaining hyperscale data centers involving millions of service hosts has been an extremely intricate task to tackle for top Internet companies.Incessant system failures cost operators countless hours of browsing through performance metrics to diagnose the underlying root cause to prevent the recurrence.Although many state-of-the-art (SOTA) methods have used time-series causal discovery to construct causal relationships among anomalous metrics, they only focus on homogeneous service-level performance metrics and fail to yield useful insights on heterogeneous host-level metrics.To address the challenge, this study presents FaultInsight, a highly interpretable deep causal host fault diagnosing framework that offers diagnostic insights from various perspectives to reduce human effort in troubleshooting.We evaluate FaultInsight using dozens of incidents collected from our production environment.FaultInsight provides markedly better root cause identification accuracy than SOTA baselines in our incident dataset.It also shows outstanding advantages in terms of deployability in real production systems.Our engineers are deeply impressed by FaultInsight's ability to interpret incidents from multiple perspectives, helping them quickly understand the mechanism behind the faults.
Tingzhu Bi, Yang Zhang 0103, Yicheng Pan 0002, Yu Zhang 0209, Meng Ma 0001, Xinrui Jiang 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003
KDD3
2024 EffCause: Discover Dynamic Causal Relationships Efficiently from Time-Series
abstract
Since the proposal of Granger causality, many researchers have followed the idea and developed extensions to the original algorithm. The classic Granger causality test aims to detect the existence of the static causal relationship. Notably, a fundamental assumption underlying most previous studies is the stationarity of causality, which requires the causality between variables to keep stable. However, this study argues that it is easy to break in real-world scenarios. Fortunately, our paper presents an essential observation: if we consider a sufficiently short window when discovering the rapidly changing causalities, they will keep approximately static and thus can be detected using the static way correctly. In light of this, we develop EffCause, bringing dynamics into classic Granger causality. Specifically, to efficiently examine the causalities on different sliding window lengths, we design two optimization schemes in EffCause and demonstrate the advantage of EffCause through extensive experiments on both simulated and real-world datasets. The results validate that EffCause achieves state-of-the-art accuracy in continuous causal discovery tasks while achieving faster computation. Case studies from cloud system failure analysis and traffic flow monitoring show that EffCause effectively helps us understand real-world time-series data and solve practical problems.
Yicheng Pan 0002, Yifan Zhang 0029, Xinrui Jiang 0001, Meng Ma 0001, Ping Wang 0003
ACM Trans. Knowl. Discov. Data1
2023 Look Deep into the Microservice System Anomaly through Very Sparse Logs
abstract
Intensive monitoring and anomaly diagnosis have become a knotty problem for modern microservice architecture due to the dynamics of service dependency. While most previous studies rely heavily on ample monitoring metrics, we raise a fundamental but always neglected issue - the diagnostic metric integrity problem. This paper solves the problem by proposing MicroCU – a novel approach to diagnose microservice systems using very sparse API logs. We design a structure named dynamic causal curves to portray time-varying service dependencies and a temporal dynamics discovery algorithm based on Granger causal intervals. Our algorithm generates a smoother space of causal curves and designs the concept of causal unimodalization to calibrate the causality infidelities brought by missing metrics. Finally, a path search algorithm on dynamic causality graphs is proposed to pinpoint the root cause. Experiments on commercial system cases show that MicroCU outperforms many state-of-the-art approaches and reflects the superiorities of causal unimodalization to raw metric imputation.
Xinrui Jiang 0001, Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003
WWW2
2023 DyCause: Crowdsourcing to Diagnose Microservice Kernel Failure
abstract
Today many web applications in the cloud (apps) are built based on microservices. However, as the anomaly propagates in a highly dynamic and complex way, troubleshooting them becomes full of challenges. Existing diagnostic methods are mostly designed based on monitoring metrics retrieved from the microservice system kernel. Therefore, application owners and even site reliability engineers (SREs) cannot effectively resort to those methods when the microservice systems lack such a comprehensive monitoring infrastructure. In this article, we develop DyCause, a crowdsourcing solution to the asymmetric diagnostic information problem. Our solution collects the operational status of kernel services collaboratively from the user space and initiates diagnosis on demand. Without the requirement of any architectural or functional infrastructure, it is both fast and lightweight to deploy DyCause in a microservice system. In order to discover the fine-grained dynamic causalities between services during the anomaly, we also design an efficient algorithm based on statistical analysis. Based on this algorithm, we can also analyze the anomaly propagation paths within the microservice system and generate a better interpretable diagnosis. In our evaluation, we test DyCause in a controlled simulation environment and a real-world cloud system. Our results have shown that DyCause has the best accuracy and efficiency among several state-of-the-art methods and is more robust in terms of parameters.
Yicheng Pan 0002, Meng Ma 0001, Xinrui Jiang 0001, Ping Wang 0003
IEEE Trans. Dependable Secur. Comput.1
2022 Abnormal Situation Simulation and Dynamic Causality Discovery in Urban Traffic Networks (Short Paper)
Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003
COSIT2
2022 When Dynamic Causality Comes to Graph-Temporal Neural Network
abstract
Spatial-temporal data forecasting is a core task in many applications, and traffic forecasting is a typical example. Researchers have proposed various methods to explore spatial and temporal characteristics to improve forecasting accuracy, including the recently emerging graph convolution networks. However, most of them only consider the road network's prior knowledge or the graph's static characteristics and thus ignore the dynamic and deeper information hidden in the data. This paper presents a novel module based on dynamic causality analysis and graph convolution to integrate statistical theories and deep learning for better capturing spatial dependencies. Then we apply the module to two specific models. In each model, we introduce the causality adjacency matrix computed by the proposed algorithm into the conventional graph convolution network to reveal the dynamic correlations between nodes in the road network. The temporal neural network is then applied to extract temporal correlations. Extensive experiments demonstrate the superiority of our method, which achieves state-of-the-art prediction accuarcy on two public transportation data sets.
Yicheng Pan 0002, Meng Ma 0001, Ping Wang 0003
IJCNN2
2022 VECROsim: A Versatile Metric-oriented Microservice Fault Simulation System (Tools and Artifact Track)
abstract
Automated fault diagnosis of microservice systems has been a hot topic in recent years. As most incidents in real commercial cloud systems are not publicly available, we have witnessed researchers putting considerable effort into developing various experimental systems. However, previous tools cannot quickly refactor their functionality, scale the architecture, and customize fault characteristics. Given this, we develop VECROsim, a versatile metric-oriented microservice fault simulation system, and release the VECROsim benchmark dataset. VECROsim works delicately as a highly-customizable toolkit to generate abnormal performance metrics datasets of microservice systems on demand and automatically. Validation of representative services from the benchmark dataset confirms the capability of VECROsim to generate realistic performance metrics for diverse real-world systems. Our case studies on root cause analysis and dynamic correlation discovery demonstrated the superiority of VECROsim. We also witnessed that the VECROsim dataset brings new research challenges to state-of-the-art fault diagnosis schemes. VECROsim concretely supports microservice developers from the industry, as well as academic researchers working on fault diagnosis or broader research topics in many ways.
Tingzhu Bi, Yicheng Pan 0002, Xinrui Jiang 0001, Meng Ma 0001, Ping Wang 0003
ISSRE2
2021 Faster, deeper, easier: crowdsourcing diagnosis of microservice kernel failure from user space
abstract
With the widespread use of cloud-native architecture, increasing web applications (apps) choose to build on microservices. Simultaneously, troubleshooting becomes full of challenges owing to the high dynamics and complexity of anomaly propagation. Existing diagnostic methods rely heavily on monitoring metrics collected from the kernel side of microservice systems. Without a comprehensive monitoring infrastructure, application owners and even cloud operators cannot resort to these kernel-space solutions. This paper summarizes several insights on operating a top commercial cloud platform. Then, for the first time, we put forward the idea of user-space diagnosis for microservice kernel failures. To this end, we develop a crowdsourcing solution - DyCause, to resolve the asymmetric diagnostic information problem. DyCause deploys on the application side in a distributed manner. Through lightweight API log sharing, apps collect the operational status of kernel services collaboratively and initiate diagnosis on demand. Deploying DyCause is fast and lightweight as we do not have any architectural and functional requirements for the kernel. To reveal more accurate correlations from asymmetric diagnostic information, we design a novel statistical algorithm that can efficiently discover the time-varying causalities between services. This algorithm also helps us build the temporal order of the anomaly propagation. Therefore, by using DyCause, we can obtain more in-depth and interpretable diagnostic clues with limited indicators. We apply and evaluate DyCause on both a simulated test-bed and a real-world cloud system. Experimental results verify that DyCause running in the user-space outperforms several state-of-the-art algorithms running in the kernel on accuracy. Besides, DyCause shows superior advantages in terms of algorithmic efficiency and data sensitivity. Simply put, DyCause produces a significantly better result than other baselines when analyzing much fewer or sparser metrics. To conclude, DyCause is faster to act, deeper in analysis, and easier to deploy.
Yicheng Pan 0002, Meng Ma 0001, Xinrui Jiang 0001, Ping Wang 0003
ISSTA1