VLDB 2026 Research / reviewers in the wild / expert
Tingzhu Bi
dblp:337/2912
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-0366-0410ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAVIAR: Disentangling Root Causes with an ICA-based VAE for Large-Scale Microservice SystemsabstractMicroservice architectures in modern software engineering generate vast quantities of heterogeneous metrics, making fault diagnosis notoriously difficult. Conventional root cause analysis (RCA) methods often struggle with high-dimensional, diverse data where only a small subset of metrics may truly drive the observed failures. In this paper, we propose CAVIAR (Causality-based Analysis via VAE and ICA for Anomaly Root-cause), a two-phase framework for interpretable RCA in large-scale microservice systems. First, we train a variational autoencoder (VAE) enhanced with Independent Component Analysis (ICA) principles to learn a low-dimensional, disentangled representation of normal microservice operation. By enforcing independence among latent variables, we discover semantically coherent factors, such as specific service loads or network-level conditions. Second, when a fault occurs, we treat anomalies as external interventions on some latent factor and optimize an interventional matrix to identify the culprit dimension. This factor is then mapped back to the original metrics for actionable diagnostics. Xinrui Jiang 0001, Tingzhu Bi, Meng Ma 0001, Ping Wang 0003 |
KDD (1) | 2 |
| 2025 | UnCLe: Towards Scalable Dynamic Causal Discovery in Non-linear Temporal SystemsabstractUncovering cause-effect relationships from observational time series is fundamental to understanding complex systems. While many methods infer static causal graphs, real-world systems often exhibit *dynamic causality*—where relationships evolve over time. Accurately capturing these temporal dynamics requires time-resolved causal graphs. We propose UnCLe, a novel deep learning method for scalable dynamic causal discovery. UnCLe employs a pair of Uncoupler and Recoupler networks to disentangle input time series into semantic representations and learns inter-variable dependencies via auto-regressive Dependency Matrices. It estimates dynamic causal influences by analyzing datapoint-wise prediction errors induced by temporal perturbations. Extensive experiments demonstrate that UnCLe not only outperforms state-of-the-art baselines on static causal discovery benchmarks but, more importantly, exhibits a unique capability to accurately capture and represent evolving temporal causality in both synthetic and real-world dynamic systems (e.g., human motion). UnCLe offers a promising approach for revealing the underlying, time-varying mechanisms of complex phenomena. Tingzhu Bi, Yicheng Pan 0002, Xinrui Jiang 0001, Huize Sun, Meng Ma 0001, Ping Wang 0003 |
NeurIPS | 1 |
| 2024 | G-Cause: Parameter-free Global Diagnosis for Hyperscale Web Service InfrastructuresabstractHyperscale web service infrastructures are becoming increasingly complex and facing a variety of threats, raising the demand for more sophisticated automated operations and diagnosis solutions. Existing anomaly root cause localization approaches often focus on Service-level components without drilling down to the lower-level resources where services are deployed, hindering the implementation of fine-grained failure fix measures. This paper introduces a challenging task called global diagnosis and addresses it by proposing a technique called G-Cause, which is applicable to both Service-level and host-level root cause analysis scenarios. G-Cause builds a highly adaptive diagnostic framework based on the frequency-domain and time-domain characteristics of monitoring metrics, allowing it to handle global diagnosis requirements from app to host with minimal parameter adjustments. We deploy and validate our approach in two typical scenarios: homogeneous metric diagnosis from app to microservice, and heterogeneous metric diagnosis for various host resources. The results demonstrate that G-Cause outperforms state-of-the-art diagnosis algorithms while providing strong interpretability. Our approach helps operators understand the core mechanism of anomaly propagation and adjust their management strategies more effectively. With these strengths, G-Cause successfully services our global product operations and also makes an impressive contribution in many other workflows. Xinrui Jiang 0001, Yang Zhang 0103, Tingzhu Bi, Xiangzhuang Shen, Yu Zhang 0209, Yicheng Pan 0002, Meng Ma 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003 |
ICWS | 3 |
| 2024 | FaultInsight: Interpreting Hyperscale Data Center Host FaultsabstractOperating and maintaining hyperscale data centers involving millions of service hosts has been an extremely intricate task to tackle for top Internet companies.Incessant system failures cost operators countless hours of browsing through performance metrics to diagnose the underlying root cause to prevent the recurrence.Although many state-of-the-art (SOTA) methods have used time-series causal discovery to construct causal relationships among anomalous metrics, they only focus on homogeneous service-level performance metrics and fail to yield useful insights on heterogeneous host-level metrics.To address the challenge, this study presents FaultInsight, a highly interpretable deep causal host fault diagnosing framework that offers diagnostic insights from various perspectives to reduce human effort in troubleshooting.We evaluate FaultInsight using dozens of incidents collected from our production environment.FaultInsight provides markedly better root cause identification accuracy than SOTA baselines in our incident dataset.It also shows outstanding advantages in terms of deployability in real production systems.Our engineers are deeply impressed by FaultInsight's ability to interpret incidents from multiple perspectives, helping them quickly understand the mechanism behind the faults. Tingzhu Bi, Yang Zhang 0103, Yicheng Pan 0002, Yu Zhang 0209, Meng Ma 0001, Xinrui Jiang 0001, Linlin Han, Feng Wang 0054, Ping Wang 0003 |
KDD | 1 |
| 2022 | VECROsim: A Versatile Metric-oriented Microservice Fault Simulation System (Tools and Artifact Track)abstractAutomated fault diagnosis of microservice systems has been a hot topic in recent years. As most incidents in real commercial cloud systems are not publicly available, we have witnessed researchers putting considerable effort into developing various experimental systems. However, previous tools cannot quickly refactor their functionality, scale the architecture, and customize fault characteristics. Given this, we develop VECROsim, a versatile metric-oriented microservice fault simulation system, and release the VECROsim benchmark dataset. VECROsim works delicately as a highly-customizable toolkit to generate abnormal performance metrics datasets of microservice systems on demand and automatically. Validation of representative services from the benchmark dataset confirms the capability of VECROsim to generate realistic performance metrics for diverse real-world systems. Our case studies on root cause analysis and dynamic correlation discovery demonstrated the superiority of VECROsim. We also witnessed that the VECROsim dataset brings new research challenges to state-of-the-art fault diagnosis schemes. VECROsim concretely supports microservice developers from the industry, as well as academic researchers working on fault diagnosis or broader research topics in many ways. Tingzhu Bi, Yicheng Pan 0002, Xinrui Jiang 0001, Meng Ma 0001, Ping Wang 0003 |
ISSRE | 1 |