Yaxiao Li

dblp:193/7431 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0009-5259-7260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DistriAD: Distributed Anomaly Detection for Large-Scale Microservice Systems
abstract
Microservice architecture is used by leading companies to develop their large-scale software systems. These systems comprise numerous nodes, diverse service types and instances, and substantial volumes of data. Current research usually requires a central node to collect massive data from the system to build an anomaly detection model, encountering two significant limitations: 1) Most research trains a model for the entire system, ignoring the unique characteristics of individual nodes. Additionally, processing vast system-wide data in a single node imposes significant resource demands. 2) Microservice systems change frequently, and the historical data distribution differs significantly from the real data distribution, resulting in concept drift. Thus, we proposes DistriAD, a distributed anomaly detection method specifically designed for large-scale microservice systems. DistriAD involves a lightweight anomaly detection model deployed on each distributed node for precise anomaly detection, thus enhancing its accuracy. Furthermore, DistriAD utilizes a federated learning framework and a continuous updating method incorporating human feedback to update model parameters and address concept drift. Experimental validation on public datasets, e.g., TrainTicket-based and GAIA, and a proprietary test system dataset demonstrate that DistriAD outperforms baseline methods, improving F1-score up to 39.2 %. We believe that this work can provide insights into distributed anomaly detection in large-scale microservice systems, thereby improving their performance.
Yaxiao Li, Qingshan Li, Chenxi Zhang 0003, Lu Wang 0014, Chenyi Wang 0001, Zhongliang Bai, Haixing Luo, Tianyuan Gao, Lingfeng Pan
ICWS1
2025 Dynamic Microservice Resource Optimization Management Based on MAPE Loop
abstract
Microservice resource management aims to ensure stable service instance loads and improve overall resource utilization through load balancing, elastic scaling, and container orchestration while meeting system service quality requirements.Existing research often focuses on localized solutions, addressing only single aspects of elastic scaling or load balancing without recognizing the systemic nature of microservice resource management.The complexity and dynamism of service dependencies make it challenging to quantify interactions between resource strategies and service loads.Additionally, microservice systems experience dynamic load variations influenced by user behavior, business activities, and external factors.The dynamic nature of load changes and the selection of load features significantly increase the difficulty of resource forecasting, further complicating microservice resource management.To address this problem, this paper proposes MDRM (MAPEbased Dynamic Resource Management), a dynamic optimization method that integrates load balancing and elastic scaling to overcome the limitations of isolated strategies.MDRM models system load based on business characteristics and service invocation relationships, accurately capturing dynamic variations.A composite model, combining parallel multi-layer CNNs and LSTMs, extracts spatiotemporal microservice features, enhancing resource forecasting accuracy.Additionally, MDRM formulates a comprehensive load balancing optimization function that synergizes with resource utilization and service response time objectives to generate optimal management strategies.Experimental results demonstrate that, compared to default resource management strategies in Docker Swarm and Kubernetes, MDRM significantly improves system throughput (approximately 1000 RPS) and reduces response time (approximately 30-40 ms), proving its effectiveness.
Lu Wang 0014, Xu Fan 0008, Yaxiao Li, Quanwei Du, Jialuo She, Qingshan Li
Internetware3
2025 Hypergraph Neural Network-based Multi-Granular Root Cause Localization for Microservice Systems
abstract
Modern enterprises are increasingly adopting microservice architectures to enhance system flexibility and scalability. However, in the face of ever-changing business requirements, the relationships between system components have become increasingly complex, resulting in significant challenges in maintaining system robustness. In recent years, multimodal data-driven approaches based on graph neural networks have emerged as a predominant solution for root cause localization in microservice systems. Our detailed analysis of architectural characteristics and existing research reveals two critical limitations. First, simple graph is insufficient to represent the one-to-many relationships inherent in microservice component interactions, such as deployment, subordinate, and dependency. Second, the current multimodal data-based method has difficulty in performing localization on faults occurring on hosts, services, and instances at the same time.To address these challenges, we propose HyperRCA, a novel multi-granular root cause analysis approach based on hypergraph neural networks. Our approach models system states during faults via a hypergraph with instances as graph nodes, explicitly capturing heterogeneous relationships through three innovative hyperedge designs: deployment hyperedges for infrastructure relationships, subordinate hyperedges for service hierarchies, and dependency hyperedges for inter-component interactions. We used hypergraph neural networks and multi-layer perceptrons to train a root cause localization model based on hyperedge features to achieve multi-granularity root cause localization. Experimental evaluations demonstrate significant performance improvements over state-of-the-art approaches. HyperRCA achieves a maximum HR@5 improvement of 112.62% on single-granularity datasets and 466.43% in multi-granularity scenarios.
Yaxiao Li, Lu Wang 0014, Chenxi Zhang 0003, Qingshan Li, Siming Rong, Baiyang Wen, Quanwei Du, KeYang Li, Lingfeng Pan, Mingxuan Hui
ASE1
2023 An Efficient Load Prediction-Driven Scheduling Strategy Model in Container Cloud
abstract
The rise of containerization has led to the development of container cloud technology, which offers container deployment and management services. However, scheduling a large number of containers efficiently remains a significant challenge for container cloud service platforms. Traditional load prediction methods and scheduling algorithms do not fully consider interdependencies between containers or fine‐grained resource scheduling, leading to poor resource utilization and scheduling efficiency. To address these challenges, this paper proposes a new load prediction model CNN‐BiGRU‐Attention and a container scheduling strategy based on load prediction. The prediction model CNN and BiGRU focus on the local features of load data and long sequence dependencies, respectively, as well as introduce the attention mechanism to make the model more easily capture the features of long distance dependencies in the sequence. A container scheduling strategy based on load prediction is also designed, which first uses the load prediction model to predict the load state and then generates a scheduling strategy based on the load prediction value to determine the change of the number of container replicas in a fine‐grained manner based on the load prediction value in the next time window, while the established domain‐based container selection method is employed to facilitate the coarse‐grained online migration of containers. Experiments conducted using public datasets and open‐source simulation platforms demonstrate that the proposed approach achieves a 37.4% improvement in container load prediction accuracy and a 21.7% improvement in container scheduling efficiency compared to traditional methods. These results highlight the effectiveness of the proposed approach in addressing the challenges faced by container cloud service platforms.
Lu Wang 0014, Shuaidong Guo, Pengli Zhang, Haodong Yue, Yaxiao Li, Chenyi Wang 0001, Zhuang Cao
Int. J. Intell. Syst.5
2016 Single image super-resolution reconstruction based on genetic algorithm and regularization prior model
Yangyang Li 0001, Yang Wang 0075, Yaxiao Li, Licheng Jiao, Xiangrong Zhang, Rustam Stolkin
Inf. Sci.3