Tiangang Li

dblp:292/4417 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0003-7582-8128ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CausalLog: Log parsing using LLMs with causal intervention for bias mitigation
Tiangang Li
Inf. Process. Manag.3
2026 DCGFI: Robust and Interpretable Failure Identification Using Dual Causal Graph for Microservices on Multimodal Observability Data
abstract
Accurate failure identification is crucial for ensuring the reliability of microservice systems. However, existing failure identification approaches are usually based on data-driven correlation learning strategies. These approaches focus on statistical correlation in observability data and ignore the causal mechanism in microservice systems, which leads to insufficient robustness and interpretability. In this study, we propose DCGFI, a robust and interpretable failure identification approach based on multimodal observability data, which uses dual causal graphs to integrate causal modeling, data-driven strategy and knowledge driven strategy to achieve effective failure identification for microservice systems. DCGFI first constructs data causal graphs and knowledge causal graphs based on multimodal observability data. Then, DCGFI uses a causal graph convolutional network to learn data causal graphs and update knowledge causal graphs. Finally, DCGFI combines data causal graphs with knowledge causal graphs to jointly optimize the failure identification process. Experimental results on two datasets demonstrate that DCGFI outperforms all baselines on both Macro-F1 and Micro-F1 and exhibits good robustness across different datasets.
Xiangbo Tian, Shi Ying 0001, Tiangang Li, Chuan Shi 0001, Ding Xiao
IEEE Trans. Dependable Secur. Comput.3
2026 DALAD: Unsupervised Detection of Global and Local Anomalies in Microservice Systems
abstract
Accurate anomaly detection is crucial for the reliability of microservice systems. However, most existing anomaly detection approaches only detect deviations from the global expected patterns, while overlooking anomalies that conform to the global expected patterns but deviate from the local expected patterns. In this paper, we propose DALAD, a novel distribution-adversarial-learning-based anomaly detection approach for microservice systems, which jointly learns the normal and anomalous system pattern distributions to effectively detect both global and local anomalies. In detail, DALAD first designs an adversarial data generation strategy to automatically generate anomalous traces at a low cost. Then, Distribution-Adversarial-Learning Trace Representation is designed to jointly learn the multivariate-Gaussian-distribution-based vector representations of normal and anomalous traces, which can reflect the difference between traces in a more fine-grained manner. Finally, it further models the normal and anomalous system pattern distributions from these vector representations, and detects anomalies by comparing the likelihoods of traces under these distributions. Experimental results on two datasets show that DALAD achieves the best anomaly detection performance while maintaining a competitive computational cost.
Xiangbo Tian, Shi Ying 0001, Tiangang Li
IEEE Trans. Serv. Comput.3
2025 DebiasParser: Debiasing LLM-Based Log Parsing via Front-Door Adjustment
Tiangang Li
ICIC (8)3
2025 TraceDAE: Trace-Based Anomaly Detection in Microservice Systems via Dual Autoencoder
abstract
Micro-service systems have become a popular architecture for modern web applications owing to their scalability, modularity, and maintainability. However, with the increasing complexity and size of these systems, anomaly detection emerges as a critical task. In this paper, we introduce TraceDAE, a trace-based anomaly detection approach in micro-service systems. The approach initially constructs a Service Trace Graph (STG) to depict service invocation relationships and performance metrics, subsequently introducing a dual autoencoder framework. In this framework, the structure autoencoder employs Graph Attention Networks (GAT) to analyze the structure, while the attribute autoencoder leverages the Long Short-Term Memory Network (LSTM) for processing time series data. This approach is capable of effectively identifying Service Response Abnormal and Service Invocation Abnormal. Moreover, the final experimental results on datasets show that TraceDAE is an efficient anomaly detection approach which outperforms the SOTA trace-based anomaly detection methods with f1-scores of 0.970 and 0.925, respectively.
Shi Ying 0001, Tiangang Li, Xiangbo Tian
IEEE Trans. Netw. Serv. Manag.3
2025 ASTRA: Adversarial Sim-to-Real Transfer Reinforcement Learning for Autoscaling in Cloud Systems
abstract
With the widespread adoption of cloud computing, autoscaling has become crucial for efficient resource management and stable service provision in cloud systems. In recent years, autoscaling methods based on deep reinforcement learning (DRL) have gained significant attention due to their outstanding adaptability and flexibility. However, training DRL-based autoscaler requires interactions with real cloud systems, incurring high interaction costs, low data collection efficiency, and potential operational impacts. To address these challenges, we propose ASTRA, a sim-to-real transfer reinforcement learning framework for autoscaling. ASTRA constructs a cloud system simulation environment based on a performance estimation model, enabling low-cost and high-efficiency training sample collection for policy learning. The learned policy is subsequently transferred to the real systems for scaling decisions. To address performance modeling inaccuracies caused by dynamic cloud state changes, we propose a performance modeling method based on hybrid attentive state space model. By incorporating state space model, it captures system dynamics and state evolution, effectively reducing simulation errors. Furthermore, to mitigate the performance degradation of the transferred policy due to the distribution shift, we propose an autoscaling method based on adversarial soft actor-critic. By introducing adversarial policy training with gradient regularization based on state perturbations, it significantly improves transferred policy performance. The results in the real system demonstrate that ASTRA achieves optimal overall performance in environment modeling, policy transfer and real-world autoscaling. Specifically, ASTRA outperforms all baselines in terms of instance number, response time, SLO violation rate, and CPU utilization under different workload patterns. More importantly, under limited interaction costs, ASTRA achieves a 616.94× improvement in interaction sample collection rate compared to direct online training method.
Tiangang Li, Shi Ying 0001, Xiangbo Tian
IEEE Trans. Software Eng.1
2024 Batch Jobs Load Balancing Scheduling in Cloud Computing Using Distributional Reinforcement Learning
abstract
In cloud computing, how to reasonably allocate computing resources for batch jobs to ensure the load balance of dynamic clusters and meet user requests is an important and challenging task. Most existing studies are based on deep Q network, which utilizes neural networks to estimate the expected value of cumulative return in the scheduling process. The value-based DQN algorithms ignore the complete information contained in the value distribution and lack strong adaptability to time-varying batch jobs and dynamic cluster resources. Therefore, to capture the inherent stochasticity of the scheduling process caused by environmental stochasticity, we utilize Distributional Reinforcement Learning to model the value distribution of the cumulative return. Specifically, we formalize the load balancing scheduling as a multi-objective optimization problem and construct a Distributional Reinforcement Learning model. Then we introduce quantile regression to learn the value distribution of the cumulative return during scheduling and propose a dynamic load balancing scheduling algorithm based on Distributional Reinforcement Learning. In addition, we develop a cluster environment for real-time processing of batch jobs to simulate the arrival of batch jobs and train the Distributional Reinforcement Learning-based scheduling agent. We conduct empirical experiments and detailed analysis by using the real Alibaba Cluster cluster traces v2018 and v2020. The results show that compared to the baseline algorithms, the proposed algorithm performs better in terms of cluster load balancing, success rate of instance creation and average completion time of the tasks. The experimental results on different trace datasets also indicate that the propsoed algorithm exhibits excellent scalability.
Tiangang Li, Shi Ying 0001, Yishi Zhao, Jianga Shang
IEEE Trans. Parallel Distributed Syst.1
2024 iTCRL: Causal-Intervention-Based Trace Contrastive Representation Learning for Microservice Systems
abstract
Nowadays, microservice architecture has become mainstream way of cloud applications delivery. Distributed tracing is crucial to preserve the observability of microservice systems. However, existing trace representation approaches only concentrate on operations, relationships and metrics related to service invocations. They ignore service events that denotes meaningful, singular point in time during the service's duration. In this paper, we propose iTCRL, a novel trace contrastive representation learning approach based on causal intervention. This approach first constructs a unified graph representation for each trace to describe the runtime status of service events in traces and the complex relationships between them. Then, Causal-intervention-based Trace Contrastive Learning is proposed, which learns trace representations from causal perspective based on the unified graph representations of traces. It uses causal intervention to generate contrastive views, heterogeneous graph neural network-based trace encoder to learn trace representations, and direct causal effect to guide the training of trace encoder. Experimental results on three datasets show that iTCRL outperforms all baselines in terms of trace classification, trace anomaly detection, trace sampling and noise robustness, and also validate the contribution of Causal-intervention-based Trace Contrastive Learning.
Xiangbo Tian, Shi Ying 0001, Tiangang Li, Mengting Yuan 0001, Ruijin Wang, Yishi Zhao, Jianga Shang
IEEE Trans. Software Eng.3