Yuqing Wang 0002

dblp:60/5086-2 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-0175-005XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AnoMod: A Dataset for Anomaly Detection and Root Cause Analysis in Microservice System
abstract
Microservice systems (MSS) have become a predominant architectural style for cloud services. Yet the community still lacks high-quality, publicly available datasets for anomaly detection (AD) and root cause analysis (RCA) in MSS. Most benchmarks emphasize performance-related faults and provide only one or two monitoring modalities, limiting research on broader failure modes and cross-modal methods. To address these gaps, we introduce a new multimodal anomaly dataset built on two open-source microservice systems: SocialNetwork and TrainTicket. We design and inject four categories of anomalies (Ano): performance-level, service-level, database-level, and code-level, to emulate realistic anomaly modes. For each scenario, we collect five modalities (Mod): logs, metrics, distributed traces, API responses, and code coverage reports, offering a richer, end-to-end view of system state and inter-service interactions. We name our dataset, reflecting its unique properties, as AnoMod. This dataset enables (1) evaluation of cross-modal anomaly detection and fusion/ablation strategies, and (2) fine-grained RCA studies across service and code regions, supporting end-to-end troubleshooting pipelines that jointly consider detection and localization.
Ke Ping, Hamza Bin Mazhar, Yuqing Wang 0002, Mika Mäntylä
MSR3
2026 Monitoring Data for Anomaly Detection in Cloud-Based Systems: A Systematic Mapping Study
abstract
Context : Anomaly detection is crucial for maintaining cloud-based software systems, as it enables early identification and resolution of unexpected failures. Given rapid and significant advances in the anomaly detection domain and the complexity of its industrial implementation, an overview of techniques that utilize real-world operational data is needed. Aim : This study aims to complement existing research with an extensive catalog of the techniques and monitoring data used for detecting anomalies affecting the performance or reliability of cloud-based software systems that have been developed and/or evaluated in a real-world context. Method : We perform a systematic mapping study to examine the literature on anomaly detection in cloud-based systems, particularly focusing on the usage of real-world monitoring data, with the aim of identifying key data categories, tools, data preprocessing, and anomaly detection techniques. Results : Based on a review of 104 papers, we categorize monitoring data by structure, types, and origins and the tools used for data collection and processing. We offer a comprehensive overview of data preprocessing and anomaly detection techniques mapped to different data categories. Our findings highlight practical challenges and considerations in applying these techniques in real-world cloud environments. Conclusion : The findings help practitioners and researchers identify relevant data categories and select appropriate data preprocessing and anomaly detection techniques for their specific operational environments, which is important for improving the reliability and performance of cloud-based systems.
Adha Hrusto, Nauman Bin Ali, Emelie Engström, Yuqing Wang 0002
ACM Trans. Softw. Eng. Methodol.4
2025 Cross-System Software Log-based Anomaly Detection Using Meta-Learning
abstract
Modern software systems produce vast amounts of logs, serving as an essential resource for anomaly detection. Artificial Intelligence for IT Operations (AIOps) tools have been developed to automate the process of log-based anomaly detection for software systems. Three practical challenges are widely recognized in this field: data labeling costs, evolving logs in dynamic systems, and adaptability across different systems. In this paper, we propose CroSysLog, an AIOps tool for log-event level anomaly detection, considering these challenges. Following prior approaches, CroSysLog uses a neural representation approach to gain a nuanced understanding of logs and generate representations for individual log events accordingly. CroSysLog can be trained on source systems with sufficient labeled logs from open datasets to achieve robustness, and then efficiently adapt to target systems with a few labeled log events for effective anomaly detection. We evaluate CroSysLog using open datasets of four large-scale distributed supercomputing systems: BGL, Thunderbird, Liberty, and Spirit. We used random log splits, maintaining the chronological order of consecutive log events, from these systems to train and evaluate CroSysLog. These splits were widely distributed across a one/two-year span of each system's log collection duration, capturing the evolving nature of the logs in each system. Our results show that, after training CroSysLog on Liberty and BGL as source systems, CroSysLog can efficiently adapt to target systems Thunderbird and Spirit using a few labeled log events from each target system, effectively performing anomaly detection for these target systems. The results demonstrate that CroSysLog is a practical, scalable, and adaptable tool for log-event level anomaly detection in operational and maintenance contexts of software systems.
Yuqing Wang 0002, Mika Mäntylä, Jesse Nyyssölä, Ke Ping
SANER1
2024 A Dataset of Microservices-based Open-Source Projects
abstract
Researchers in the microservices community often resort to demonstrating the impact of their proposed advancements on custom-made microservices projects. This is a possible source of bias that can reduce the trustworthiness of the results. Moreover, it is hard to compare advances in small projects, often developed due to lack of time. It is common across disciplines to recognize benchmarks that mitigate bias and unify the advancements' impact. To facilitate the identification of available open-source microservice projects (OSS-MS), we performed a comprehensive study to identify, curate, and catalog OSS-MS. We started with 389559 projects and filtered them down to 3804 projects that we manually labeled. After manual labeling, our dataset contains 378 projects with three or more microservices and with over 100 commits. We document the projects from many perspectives, including project size, platform, number of contributors, project purpose, and foundation support. This dataset can serve researchers as a roadmap to identify benchmarks, as our dataset can be used to answer questions such as whether the number of services impacts the issue count.
Dario Amoroso d'Aragona, Alexander Bakhtin, Xiaozhou Li 0002, Ruoyu Su, Lauren Adams, Ernesto Aponte, Francis Boyle, Patrick Boyle, Rachel Koerner, Joseph Lee, Fangchao Tian, Yuqing Wang 0002, Jesse Nyyssölä, Ernesto Quevedo Caballero, Md Shahidur Rahaman, Amr S. Abdelfattah, Mika Mäntylä, Tomás Cerný, Davide Taibi 0001
MSR12
2024 LogLead - Fast and Integrated Log Loader, Enhancer, and Anomaly Detector
abstract
This paper introduces LogLead, a tool designed for efficient log analysis benchmarking. LogLead combines three essential steps in log processing: loading, enhancing, and anomaly detection. The tool leverages Polars, a high-speed DataFrame library. We currently have Loaders for eight systems that are publicly available (HDFS, Hadoop, BGL, Thunderbird, Spirit, Liberty, TrainTicket, and GC Webshop). We have multiple enhancers with three parsers (Drain, Spell, LenMa), Bert embedding creation and other log representation techniques like bag-of-words. LogLead integrates to five supervised and four unsupervised machine learning algorithms for anomaly detection from SKLearn. By integrating diverse datasets, log representation methods and anomaly detectors, LogLead facilitates comprehensive benchmarking in log analysis research. We show that log loading from raw file to dataframe is over 10x faster with LogLead compared to past solutions. We demonstrate roughly 2x improvement in Drain parsing speed by off-loading log message normalization to LogLead. Our brief benchmarking on HDFS indicates that log representations extending beyond the bag-of-words approach offer limited additional benefits. Tool URL: https://github.com/EvoTestOps/LogLead.
Mika Mäntylä, Yuqing Wang 0002, Jesse Nyyssölä
SANER2
2022 Test automation maturity improves product quality - Quantitative study of open source projects using continuous integration
abstract
The popularity of continuous integration (CI) is increasing as a result of market pressure to release product features or updates frequently. The ability of CI to deliver quality at speed depends on reliable test automation. In this paper, we present an empirical study to observe the effect of test automation maturity (assessed by standard best practices in the literature) on product quality, test automation effort, and release cycle in the CI context of open source projects. We run our test automation maturity survey and got responses from 37 open source java projects. We also mined software repositories of the same projects. The main results of regression analysis reveal that, higher levels of test automation maturity are positively associated with higher product quality (p-value=0.000624) and shorter release cycle (p-value=0.01891); There is no statistically significant evidence of increased test automation effort due to higher levels of test automation maturity and product quality. Thus, we conclude that, a potential benefit of improving test automation maturity (using standard best practices) is product quality improvement and release cycle acceleration in the CI context of open source projects. We encourage future research to extend our findings by adding more datasets with different programming languages and CI tools, closed source projects, and large-scale industrial projects. Our recommendation to practitioners (in the similar CI context) is to utilize standard best practices to improve test automation maturity.
Yuqing Wang 0002, Mika Mäntylä, Jouni Markkula
J. Syst. Softw.1
2022 Improving test automation maturity: A multivocal literature review
abstract
Abstract Mature test automation is key for achieving software quality at speed. In this paper, we present a multivocal literature review with the objective to survey and synthesize the guidelines given in the literature for improving test automation maturity. We selected and reviewed 81 primary studies, consisting of 26 academic literature and 55 grey literature sources. From primary studies, we extracted 26 test automation best practices (e.g., Define an effective test automation strategy, Set up good test environments, and Develop high‐quality test scripts) and collected many pieces of advice (e.g., in forms of implementation/improvement approaches, technical techniques, concepts, and experience‐based heuristics) on how to conduct these best practices. We made main observations: (1) There are only six best practices whose positive effect on maturity improvement have been evaluated by academic studies using formal empirical methods; (2) several technical related best practices in this MLR were not presented in test maturity models; (3) some best practices can be linked to success factors and maturity impediments proposed by other scholars; (4) most pieces of advice on how to conduct proposed best practices were identified from experience studies and their effectiveness need to be further evaluated with cross‐site empirical evidence using formal empirical methods; (5) in the literature, some advice on how to conduct certain best practices are conflicting, and some advice on how to conduct certain best practices still need further qualitative analysis.
Yuqing Wang 0002, Mika Mäntylä, Jouni Markkula, Päivi Raulamo-Jurvanen
Softw. Test. Verification Reliab.1
2020 Software Test Automation Maturity: A Survey of the State of the Practice
abstract
The software industry has seen an increasing interest in test automation. In this paper, we present a test automation maturity survey serving as a self-assessment for practitioners. Based on responses of 151 practitioners coming from above 101 organizations in 25 countries, we make observations regarding the state of the practice of test automation maturity: a) The level of test automation maturity in different organizations is differentiated by the practices they adopt; b) Practitioner reported the quite diverse situation with respect to different practices, e.g., 85\% practitioners agreed that their test teams have enough test automation expertise and skills, while 47\% of practitioners admitted that there is lack of guidelines on designing and executing automated tests; c) Some practices are strongly correlated and/or closely clustered; d) The percentage of automated test cases and the use of Agile and/or DevOps development models are good indicators for a higher test automation maturity level; (e) The roles of practitioners may affect response variation, e.g., QA engineers give the most optimistic answers, consultants give the most pessimistic answers. Our results give an insight into present test automation processes and practices and indicate chances for further improvement in the present industry.
Yuqing Wang 0002, Mika Mäntylä, Serge Demeyer, Kristian Wiklund, Sigrid Eldh, Tatu Kairi
ICSOFT1
2019 A Self-assessment Instrument for Assessing Test Automation Maturity
abstract
Test automation is important in the software industry but self-assessment instruments for assessing its maturity are not sufficient. The two objectives of this study are to synthesize what an organization should focus to assess its test automation; develop a self-assessment instrument (a survey) for assessing test automation maturity and scientifically evaluate it. We carried out the study in four stages. First, a literature review of 25 sources was conducted. Second, the initial instrument was developed. Third, seven experts from five companies evaluated the initial instrument. Content Validity Index and Cognitive Interview methods were used. Fourth, we revised the developed instrument. Our contributions are as follows: (a) we collected practices mapped into 15 key areas that indicate where an organization should focus to assess its test automation; (b) we developed and evaluated a self-assessment instrument for assessing test automation maturity; (c) we discuss important topics such as response bias that threatens self-assessment instruments. Our results help companies and researchers to understand and improve test automation practices and processes.
Yuqing Wang 0002, Mika Mäntylä, Sigrid Eldh, Jouni Markkula, Kristian Wiklund, Tatu Kairi, Päivi Raulamo-Jurvanen, Antti Haukinen
EASE1
2018 Test Automation Maturity Assessment
Yuqing Wang 0002
ICST1
2017 Cultural Factors Influencing International Collaborative Software Engineering Education in China
abstract
Software engineering (SE) is a rapidly developing international discipline that requires up-to-date knowledge and skills. The need for well-educated professional software engineers is increasing globally. In China, universities are opening opportunities for collaboration and building cooperative relationships with Western universities in technology fields, including SE, to offer Chinese students possibilities for international education in China instead of studying abroad. Designing high-quality SE education in international collaborative programs faces challenges introduced by cultural factors that affect learning practices. In this study, we addressed these challenges in the context of international collaborative SE education in China. In the first step, we synthesized existing knowledge of Chinese cultural factors affecting learning by conducting a systematic literature review (SLR). In the second step, we conducted interviews with SE students and teachers in a Chinese university that is preparing an international collaborative SE program, in order to see whether the identified cultural factors are valid in the current learning contexts of SE education. The results revealed that many of the identified factors are still valid, but some of them present differently in the current context because of the novelty of the SE discipline and the changing educational environment in China.
Yuqing Wang 0002, Jouni Markkula
APSEC1