Shi Ying 0001

dblp:75/3547-1 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-0471-0021ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DCGFI: Robust and Interpretable Failure Identification Using Dual Causal Graph for Microservices on Multimodal Observability Data
abstract
Accurate failure identification is crucial for ensuring the reliability of microservice systems. However, existing failure identification approaches are usually based on data-driven correlation learning strategies. These approaches focus on statistical correlation in observability data and ignore the causal mechanism in microservice systems, which leads to insufficient robustness and interpretability. In this study, we propose DCGFI, a robust and interpretable failure identification approach based on multimodal observability data, which uses dual causal graphs to integrate causal modeling, data-driven strategy and knowledge driven strategy to achieve effective failure identification for microservice systems. DCGFI first constructs data causal graphs and knowledge causal graphs based on multimodal observability data. Then, DCGFI uses a causal graph convolutional network to learn data causal graphs and update knowledge causal graphs. Finally, DCGFI combines data causal graphs with knowledge causal graphs to jointly optimize the failure identification process. Experimental results on two datasets demonstrate that DCGFI outperforms all baselines on both Macro-F1 and Micro-F1 and exhibits good robustness across different datasets.
Xiangbo Tian, Shi Ying 0001, Tiangang Li, Chuan Shi 0001, Ding Xiao
IEEE Trans. Dependable Secur. Comput.2
2026 DALAD: Unsupervised Detection of Global and Local Anomalies in Microservice Systems
abstract
Accurate anomaly detection is crucial for the reliability of microservice systems. However, most existing anomaly detection approaches only detect deviations from the global expected patterns, while overlooking anomalies that conform to the global expected patterns but deviate from the local expected patterns. In this paper, we propose DALAD, a novel distribution-adversarial-learning-based anomaly detection approach for microservice systems, which jointly learns the normal and anomalous system pattern distributions to effectively detect both global and local anomalies. In detail, DALAD first designs an adversarial data generation strategy to automatically generate anomalous traces at a low cost. Then, Distribution-Adversarial-Learning Trace Representation is designed to jointly learn the multivariate-Gaussian-distribution-based vector representations of normal and anomalous traces, which can reflect the difference between traces in a more fine-grained manner. Finally, it further models the normal and anomalous system pattern distributions from these vector representations, and detects anomalies by comparing the likelihoods of traces under these distributions. Experimental results on two datasets show that DALAD achieves the best anomaly detection performance while maintaining a competitive computational cost.
Xiangbo Tian, Shi Ying 0001, Tiangang Li
IEEE Trans. Serv. Comput.2
2025 Performance issue monitoring, identification and diagnosis of SaaS software: a survey
Rui Wang 0036, Xiangbo Tian, Shi Ying 0001
Frontiers Comput. Sci.3
2025 TraceDAE: Trace-Based Anomaly Detection in Microservice Systems via Dual Autoencoder
abstract
Micro-service systems have become a popular architecture for modern web applications owing to their scalability, modularity, and maintainability. However, with the increasing complexity and size of these systems, anomaly detection emerges as a critical task. In this paper, we introduce TraceDAE, a trace-based anomaly detection approach in micro-service systems. The approach initially constructs a Service Trace Graph (STG) to depict service invocation relationships and performance metrics, subsequently introducing a dual autoencoder framework. In this framework, the structure autoencoder employs Graph Attention Networks (GAT) to analyze the structure, while the attribute autoencoder leverages the Long Short-Term Memory Network (LSTM) for processing time series data. This approach is capable of effectively identifying Service Response Abnormal and Service Invocation Abnormal. Moreover, the final experimental results on datasets show that TraceDAE is an efficient anomaly detection approach which outperforms the SOTA trace-based anomaly detection methods with f1-scores of 0.970 and 0.925, respectively.
Shi Ying 0001, Tiangang Li, Xiangbo Tian
IEEE Trans. Netw. Serv. Manag.2
2025 ASTRA: Adversarial Sim-to-Real Transfer Reinforcement Learning for Autoscaling in Cloud Systems
abstract
With the widespread adoption of cloud computing, autoscaling has become crucial for efficient resource management and stable service provision in cloud systems. In recent years, autoscaling methods based on deep reinforcement learning (DRL) have gained significant attention due to their outstanding adaptability and flexibility. However, training DRL-based autoscaler requires interactions with real cloud systems, incurring high interaction costs, low data collection efficiency, and potential operational impacts. To address these challenges, we propose ASTRA, a sim-to-real transfer reinforcement learning framework for autoscaling. ASTRA constructs a cloud system simulation environment based on a performance estimation model, enabling low-cost and high-efficiency training sample collection for policy learning. The learned policy is subsequently transferred to the real systems for scaling decisions. To address performance modeling inaccuracies caused by dynamic cloud state changes, we propose a performance modeling method based on hybrid attentive state space model. By incorporating state space model, it captures system dynamics and state evolution, effectively reducing simulation errors. Furthermore, to mitigate the performance degradation of the transferred policy due to the distribution shift, we propose an autoscaling method based on adversarial soft actor-critic. By introducing adversarial policy training with gradient regularization based on state perturbations, it significantly improves transferred policy performance. The results in the real system demonstrate that ASTRA achieves optimal overall performance in environment modeling, policy transfer and real-world autoscaling. Specifically, ASTRA outperforms all baselines in terms of instance number, response time, SLO violation rate, and CPU utilization under different workload patterns. More importantly, under limited interaction costs, ASTRA achieves a 616.94× improvement in interaction sample collection rate compared to direct online training method.
Tiangang Li, Shi Ying 0001, Xiangbo Tian
IEEE Trans. Software Eng.2
2024 Batch Jobs Load Balancing Scheduling in Cloud Computing Using Distributional Reinforcement Learning
abstract
In cloud computing, how to reasonably allocate computing resources for batch jobs to ensure the load balance of dynamic clusters and meet user requests is an important and challenging task. Most existing studies are based on deep Q network, which utilizes neural networks to estimate the expected value of cumulative return in the scheduling process. The value-based DQN algorithms ignore the complete information contained in the value distribution and lack strong adaptability to time-varying batch jobs and dynamic cluster resources. Therefore, to capture the inherent stochasticity of the scheduling process caused by environmental stochasticity, we utilize Distributional Reinforcement Learning to model the value distribution of the cumulative return. Specifically, we formalize the load balancing scheduling as a multi-objective optimization problem and construct a Distributional Reinforcement Learning model. Then we introduce quantile regression to learn the value distribution of the cumulative return during scheduling and propose a dynamic load balancing scheduling algorithm based on Distributional Reinforcement Learning. In addition, we develop a cluster environment for real-time processing of batch jobs to simulate the arrival of batch jobs and train the Distributional Reinforcement Learning-based scheduling agent. We conduct empirical experiments and detailed analysis by using the real Alibaba Cluster cluster traces v2018 and v2020. The results show that compared to the baseline algorithms, the proposed algorithm performs better in terms of cluster load balancing, success rate of instance creation and average completion time of the tasks. The experimental results on different trace datasets also indicate that the propsoed algorithm exhibits excellent scalability.
Tiangang Li, Shi Ying 0001, Yishi Zhao, Jianga Shang
IEEE Trans. Parallel Distributed Syst.2
2024 iTCRL: Causal-Intervention-Based Trace Contrastive Representation Learning for Microservice Systems
abstract
Nowadays, microservice architecture has become mainstream way of cloud applications delivery. Distributed tracing is crucial to preserve the observability of microservice systems. However, existing trace representation approaches only concentrate on operations, relationships and metrics related to service invocations. They ignore service events that denotes meaningful, singular point in time during the service's duration. In this paper, we propose iTCRL, a novel trace contrastive representation learning approach based on causal intervention. This approach first constructs a unified graph representation for each trace to describe the runtime status of service events in traces and the complex relationships between them. Then, Causal-intervention-based Trace Contrastive Learning is proposed, which learns trace representations from causal perspective based on the unified graph representations of traces. It uses causal intervention to generate contrastive views, heterogeneous graph neural network-based trace encoder to learn trace representations, and direct causal effect to guide the training of trace encoder. Experimental results on three datasets show that iTCRL outperforms all baselines in terms of trace classification, trace anomaly detection, trace sampling and noise robustness, and also validate the contribution of Causal-intervention-based Trace Contrastive Learning.
Xiangbo Tian, Shi Ying 0001, Tiangang Li, Mengting Yuan 0001, Ruijin Wang, Yishi Zhao, Jianga Shang
IEEE Trans. Software Eng.2
2022 An empirical study of the effectiveness of IR-based bug localization for large-scale industrial projects
Wei Li 0241, Qing'an Li, Yunlong Ming, Weijiao Dai, Shi Ying 0001, Mengting Yuan 0001
Empir. Softw. Eng.5
2020 SaaS software performance issue diagnosis using independent component analysis and restricted Boltzmann machine
abstract
Summary SaaS software performance issue diagnosis aims to classify the type of the performance records. Deep classification method has gained much attention as a way to construct hierarchical representations from a small amount of labeled data. However, there are few researches on how to solve the classification problem of performance issues by using the deep classification method. In addition, shallow classification methods exist some problems, such as the training sample is large and the ability to fit complex functions is weak. In this article, we proposed a deep performance issue classification method based on Independent Component Analysis (ICA) and Restricted Boltzmann Machine (RBM). ICA is used to extract the features, after this process, the classification feature is obtained as RBM input, and the extracted information about performance issue is transformed into identifiable information for the classifier via visible structure of input; Hidden layer for RBM is built to realize the data transmission between hidden structure, keeping the key information; And the classification algorithm is implemented to solve our performance issue diagnosis problem of SaaS software. Experiments show that the performance of our approach is superior to the classical shallow classification algorithm, and it also meet the efficiency requirement.
Rui Wang 0036, Shi Ying 0001
Concurr. Comput. Pract. Exp.2
2020 Log-Based Anomaly Detection with the Improved K-Nearest Neighbor
abstract
Logs play an important role in the maintenance of large-scale systems. The number of logs which indicate normal (normal logs) differs greatly from the number of logs that indicate anomalies (abnormal logs), and the two types of logs have certain differences. To automatically obtain faults by K-Nearest Neighbor (KNN) algorithm, an outlier detection method with high accuracy, is an effective way to detect anomalies from logs. However, logs have the characteristics of large scale and very uneven samples, which will affect the results of KNN algorithm on log-based anomaly detection. Thus, we propose an improved KNN algorithm-based method which uses the existing mean-shift clustering algorithm to efficiently select the training set from massive logs. Then we assign different weights to samples with different distances, which reduces the negative effect of unbalanced distribution of the log samples on the accuracy of KNN algorithm. By comparing experiments on log sets from five supercomputers, the results show that the method we proposed can be effectively applied to log-based anomaly detection, and the accuracy, recall rate and F measure with our method are higher than those of traditional keyword search method.
Bingming Wang, Shi Ying 0001, Guoli Cheng, Rui Wang 0036, Bo Dong 0006
Int. J. Softw. Eng. Knowl. Eng.2
2020 HSACMA: a hierarchical scalable adaptive cloud monitoring architecture
Rui Wang 0036, Shi Ying 0001, Shun Jia
Softw. Qual. J.2
2019 Log Data Modeling and Acquisition in Supporting SaaS Software Performance Issue Diagnosis
abstract
Data logging is helpful for the operation and maintenance manager of SaaS-based solutions to diagnose performance issues. However, long-running SaaS software may generate huge amounts of log data which is difficult to analyze, and it lacks a systematic approach to collect the running log and lacks a unified data structure to normalize the performance-related data. All these threaten the timeliness of SaaS performance issue diagnosis. In this paper, we propose an architecture for log collection and analysis to support the assessment of performance and diagnosis of performance issues of SaaS-based application in cloud computing. The architecture has the three-tier structure and includes a pivot data model to integrate heterogeneous log. The two high-level metrics in the model of Average Response Time (ART) and Request Timeout Rate (RTR) are calculated by statistical measurement and the lower-level metrics are monitored in real-time. Operation and maintenance managers can evaluate the performance of SaaS software based on the high-level metrics, then timely locate the issues from the low-level metrics and take appropriate measures. Thereupon, this study presents the general-purpose technique for the architecture to support real-time big log data collection, access, computation, storage. The proposal has been implemented and validated in a case study.
Rui Wang 0036, Shi Ying 0001, Xiangyang Jia
Int. J. Softw. Eng. Knowl. Eng.2
2018 Cover Image Volume 48, Issue 11
abstract
The cover image, by Rui Wang and Shi Ying., is based on the Research Article SaaS Software Performance Issue Identification Using HMRF-MAP Framework, https://doi.org/10.1002/spe.2607.
Rui Wang 0036, Shi Ying 0001
Softw. Pract. Exp.2
2018 SaaS software performance issue identification using HMRF-MAP framework
abstract
Summary The performance of the software‐as‐a‐service (SaaS) software is often characterized by combinations of performance metrics monitored in a cloud computing platform. Due to the complexity of the application software and the dynamic nature of the deployment environment, manual diagnosis for performance issues based on metric data is typically expensive and laborious. In order to solve the above problems, we propose an automatic performance issue identification method. This approach constructs the hidden Markov random field maximum a posteriori (HMRF‐MAP) model based on the monitored metric values. The model calculates the current performance state of the system by analyzing the historical states of the system. In this paper, we evaluate our approach in a case study of a production system deployed on the cloud computing platform. The evaluation results show that our approach (1) has small system overhead, (2) is accurate in identifying the time frame during which a performance issue occurs, (3) is indeed useful and assists an operation and maintenance manager in recovering the service capability of SaaS software, and (4) is better than other approaches for identifying the performance issues in the system.
Rui Wang 0036, Shi Ying 0001
Softw. Pract. Exp.2
2017 Model Construction and Data Management of Running Log in Supporting SaaS Software Performance Analysis
abstract
Changes in operating environment may result in the performance degradation to a SaaS software.Analyzing running log is an efficient method to locate this problem.However, as a long-running software, SaaS may generate huge log data which is difficult to analyze, and it lacks a systematic approach to implement management of the running log.These all threaten the timeliness of SaaS performance analysis.In this paper, we define a log format to standardize multi-source heterogeneous log data and construct a log model to support SaaS software performance analysis, where the two performance metrics of average response time and request timeout rate in the model are calculated by statistical measurement.Furthermore, a log management framework is given to support real-time big log data collection, access, calculation, storage and service, and the technology implementation of the framework is also given.Finally, a case study is given to illustrate and validate the effectiveness of the approach.
Rui Wang 0036, Shi Ying 0001, Chengai Sun, Hongyan Wan, Huo-lin Zhang, Xiangyang Jia
SEKE2