Xiaosong Huang

dblp:197/3699 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-8117-7323ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UDA-RCL: Unsupervised Domain Adaptation for Microservice Root Cause Localization Utilizing Multimodal Data
abstract
Root cause Localization methods play a crucial role in ensuring the stability of large-scale microservice systems. However, existing methods either rely on unsupervised approaches with limited localization accuracy or on supervised learning that requires large volumes of historical anomaly data. Such data is often unavailable in newly deployed systems. To address this limitation, we attempt to use the labeled data from mature systems to help new systems build root cause localization models. Specifically, we propose UDA-RCL, an unsupervised domain adaptation root cause localization method using multimodal data (log, metric and trace data). UDA-RCL first incorporates an aggregation based event extraction module to standardize the format of multimodal data from different systems. Then, it utilizes a multimodal event encoder and multimodal domain adversarial adaptation module to narrow the feature distribution gap between different systems. Furthermore, taking into account the situation of sparse anomaly samples, existing methods' classifiers are either hard to transfer to other systems or struggle to capture the process of anomaly propagation, we propose a PageRank classifier module. This module employs a neural network embedded with anomaly propagation rules to output the final root cause ranking results, alleviating the issue of sparse anomaly samples. Extensive experiments have proven that our method achieves the best results in both supervised and transfer learning scenarios.
Xiaosong Huang, Yifan Wu 0002, Lingzhe Zhang, Ying Li 0012, Zhonghai Wu
IEEE Trans. Serv. Comput.1
2026 NER-AD: Noise-Robust Reconstruction Enhanced by Representation-Learning for Metric Anomaly Detection in Online Service Systems
Xiaosong Huang, Mengxi Jia, Zhonghai Wu, Ying Li 0012, Yu-an Tan 0001, Liehuang Zhu, Wanlei Zhou 0001
IEEE Trans. Serv. Comput.2
2025 AAAD: Asynchronous Inter-Variable Relationship-Aware Anomaly Detection for Multivariate Time Series
abstract
Anomaly detection in multivariate time series (MTS) plays a crucial role in various domains, particularly in multimedia. While significant progress has been made in modeling normal data patterns and detecting anomalies based on deviations, existing methods face challenges in capturing complex inter-variable relationships, particularly in the presence of asynchronous dependencies. To address the challenges of detecting anomalies in MTS with complex inter-variable relationships, we propose an asynchronous inter-variable relationship-aware anomaly detection method. This approach simultaneously extracts both synchronous and asynchronous feature pairs between variables and leverages attention mechanisms to automatically learn unified inter-variable relationships across these dependencies. Additionally, we utilize a memory network to store and dynamically update the normal patterns of inter-variable relationships, computing anomaly scores based on deviations in these relationships. Extensive experiments demonstrate that our method outperforms existing baseline approaches, achieving an average F1 score of 97.05% across five benchmark datasets. Ablation studies further validate the effectiveness of each component of our method.
Xiaosong Huang, Mengxi Jia, Lingzhe Zhang, Zhonghai Wu, Ying Li 0012
ICME2
2025 ORA: Job Runtime Prediction for High-Performance Computing Platforms Using the Online Retrieval-Augmented Language Model
abstract
Accurate job runtime prediction is critical for efficient scheduling in high-performance computing (HPC) platforms.For instance, precise predictions enable techniques such as backfilling, where small jobs are executed ahead of schedule to maximize resource utilization and enhance computational efficiency.However, existing runtime prediction methods primarily rely on job metadata (e.g., submission time, requested runtime, and required memory) while ignoring the content of job scripts, which limits their accuracy.To address this issue, we propose an Online Retrieval-Augmented Language Model (ORA) for job runtime prediction.ORA encodes both metadata and script information from historical jobs into feature vectors to form a database, enabling similarity-based retrieval to assist in predicting the runtime of new jobs.To address distribution shifts, ORA incrementally updates the database without requiring model retraining.Additionally, personalized retrieval mechanisms are employed to mitigate the * Co-corresponding author.
Yinping Ma, Xiaosong Huang, Lingzhe Zhang, Ying Li 0012
ICS3
2024 Cross Online Ride-Sharing for Multiple-Platform Cooperations in Spatial Crowdsourcing
abstract
The last few years have seen the wide applications of ride-sharing, a transportation service that allows users to share their travel routes. A typical problem for ride-sharing is to find an optimal route for each worker to serve the dynamically arriving requests with different objectives. Previous studies focus on the route planning on a single platform. However, a single platform may have an uneven distribution of supply and demand, which causes the platform to lose requests from lack of available workers. Luckily, some ride-sharing platforms provide the same service, which enables their collaborations. The inter-platform collaborations on ride-sharing can ease the worker shortages and greatly improve the service quality, but have not been studied yet. In this paper, we propose a Cross Online Ride-sharing (CORS) problem, which allows a platform to borrow the available workers from other platforms to serve its own requests. We first design two algorithms to select the optimal available worker from other platforms, ROWS and DOWS. ROWS randomly picks an available worker, while DOWS selects the optimal worker with the minimum additional travel distance calculated based on his/er predicted destination direction. Then, we design an efficient CORS framework that embeds the proposed optimal worker selection algorithms for the CORS problem. Extensive experiments on real and synthetic datasets demonstrate the effectiveness and efficiency of our algorithms.
Yurong Cheng, Zhaohe Liao, Xiaosong Huang, Yi Yang 0032, Xiangmin Zhou, Ye Yuan 0001, Guoren Wang
ICDE3
2024 OCRCL: Online Contrastive Learning for Root Cause Localization of Business Incidents
abstract
Microservices architecture has garnered extensive attention for its stability and scalability. However, in the complex and dynamic landscape of microservices systems, a incident in one service can propagate to others, resulting in significant economic losses and degraded user experiences. Therefore, the effective and precise localization of incidents in microservices systems becomes a critical concern. Previous research has leveraged runtime data (logs, metrics, call traces) and historical incident data to assist in root cause localization. However, due to the scarcity of business incidents (those causing severe impacts on business operations) and the fact that many incidents are reported by users, relevant run-time data and sufficient historical data are often unavailable, rendering previous methods impractical. In response to this challenge, we propose an online contrastive learning-based method for root cause localization of business incidents(OCRCL). We fully exploit incident tickets and the static dependency graph of services, integrating both textual semantic information and structural information from the dependency graph to discover root causes. Furthermore, we suggest that online contrastive learning can exhibit excellent performance with limited data and enable real-time model updates, making it better suited for industrial scenarios. Our approach demonstrates significant improvements over baseline methods across three real-world industrial datasets, highlighting its effectiveness in root cause localization.
Xiaosong Huang, Yifan Wu 0002, Yujin Zhao, Changlong Wu, Songlin Zhang, Ying Li 0012, Zhonghai Wu
SANER1
2024 UAC-AD: Unsupervised Adversarial Contrastive Learning for Anomaly Detection on Multi-Modal Data in Microservice Systems
abstract
To ensure the stability and reliability of microservice systems, timely and accurate anomaly detection is of utmost importance. Recently, considering the lack of labels in real-world scenarios and the collaborative and complementary relationships of multi-modal data in reflecting system anomalies, unsupervised multi-modal anomaly methods have been proposed. However, existing methods face challenges in effectively distinguishing normal hard samples (they are normal but hard to classify correctly) from anomalies. This is mainly caused by two aspects. First, the hard sample patterns are complex. Second, the convergence speed is inconsistent between hard and simple samples. To overcome these issues, we propose an unsupervised adversarial contrastive multi-modal anomaly detection method (UAC-AD). We utilize contrastive learning to help learn the complex patterns of hard samples and enlarge the distance between hard and anomaly samples. Meanwhile, the adversarial framework automatically identifies hard samples and fine-grained adjusts the training weights to each modality part of these hard samples. In this case, The hard sample problems of two aspects can be alleviated. We extensively evaluate UAC-AD on two open-source simulated datasets and a real industrial dataset from a large communication company. Extensive experimental results demonstrate the effectiveness of our approach in anomaly detection. We also release the code and dataset for replication and future research.
Xiaosong Huang, Mengxi Jia, Zhonghai Wu, Ying Li 0012
IEEE Trans. Serv. Comput.2
2023 Identifying Root-Cause Changes for User-Reported Incidents in Online Service Systems
abstract
In online service systems, a majority of incidents are caused by changes, which can influence user experience and cause huge economic loss. Experiences with a real-world, large-scale online service system show that more than half of the change-induced incidents are reported by users. Identifying root-cause changes for these incidents is challenging due to the inherent gap between user-perceived functional-level incident information and component-level change details. Inadequate causal knowledge also brings challenges. In this paper, we propose a novel causal knowledge mining based approach aiming at root-cause change identification for user-reported incidents named Raccoon. To bridge the gap between incidents and changes, it utilizes the fault tree and software product line to represent incidents and changes at the user-perceived functional level. They are also used as the backbone of causal knowledge. To overcome the lack of causal knowledge, Raccoon adopts efficient knowledge extraction and inference methods. Moreover, Raccoon provides recommendations at the software product line and change granularity to meet diverse demands of incident triage and root-cause change identification scenarios in incident management. We evaluate Raccoon on a real-world dataset collected in a large-scale online service system. The result shows that Raccoon significantly outperforms the state-of-the-art baseline approaches, which proves its effectiveness.
Yujin Zhao, Ye Tao 0011, Songlin Zhang, Changlong Wu, Xiaosong Huang, Ying Li 0012, Zhonghai Wu
ISSRE7
2023 UDA-DP: Unsupervised Domain Adaptation for Software Defect Prediction
abstract
Software defect prediction can automatically locate defective code modules to focus testing resources better. Traditional defect prediction methods mainly focus on manually designing features, which are input into machine learning classifiers to identify defective code. However, there are mainly two problems in prior works. First manually designing features is time consuming and unable to capture the semantic information of programs, which is an important capability for accurate defect prediction. Second the labeled data is limited along with severe class imbalance, affecting the performance of defect prediction.In response to the above problems, we first propose a new unsupervised domain adaptation method using pseudo labels for defect prediction(UDA-DP). Compared to manually designed features, it can automatically extract defective features from source programs to save time and contain more semantic information of programs. Moreover, unsupervised domain adaptation using pseudo labels is a kind of transfer learning, which is effective in leveraging rich information of limited data, alleviating the problem of insufficient data.Experiments with 10 open source projects from the PROMISE data set show that our proposed UDA-DP method outperforms the state-of-the-art methods for both within-project and cross-project defect predictions. Our code and data are available at https://github.com/xsarvin/UDA-DP.
Xiaosong Huang, Yifan Wu 0002, Ying Li 0012, Hao Yu 0016, Dadi Guo, Zhonghai Wu
SANER1