Gaolei Yi

dblp:283/7995 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 An Empirical Study on Machine Learning-Based Risk Prediction for Petroleum Pipelines
abstract
Effective risk prediction is essential for ensuring the safe operation of petroleum pipelines within the framework of pipeline integrity management. This study empirically investigates the application of machine learning techniques to pipeline risk prediction. By performing a correlation analysis on a risk-related dataset, the study identifies key relationships between various features and risk events, laying the groundwork for model development. Five machine learning algorithms-Decision Tree, Random Forest, Support Vector Machine, Neural Network, and Gradient Boosting Decision Tree (GBDT)-are implemented and evaluated. Experimental results indicate that the GBDT model outperforms the others, achieving an accuracy of 91%, precision of 94%, and recall of 96%. After parameter tuning, the GBDT model achieves a significantly improved accuracy of 99%. These results demonstrate that the GBDT-based model can effectively classify pipeline risk levels into high, medium-high, medium, and low categories, offering a robust and data-driven basis for risk identification and management in petroleum pipeline systems.
Haikang Gao, Ye Shang, Gaolei Yi, Yuan Zhao 0010, Zhenyu Chen 0001
QRS3
2025 Chattss: Improving Test Suite Simplification Via Large Language Models
abstract
As a critical component of software testing activities, regression testing plays an indispensable role in ensuring the correctness of software systems after changes. With the increasing scale and complexity of modern software, a pressing challenge arises: how to efficiently select the most effective test cases from existing test suites for regression testing, thereby reducing the associated cost. Although numerous methods have been proposed for test suite reduction, most of them rely on the assumption that test cases are independent of each other. In this paper, we present ChatTSS, a novel test case simplification approach powered by LLM. Unlike conventional test suite reduction that only shrinks the size of the test suite without altering individual test cases, ChatTSSleverages the program analysis capabilities of LLM to decompose test cases into fine-grained test atoms. It then applies appropriate reduction algorithms to perform more precise and effective test suite simplification. We conducted experiments on seven open-source projects, comprising over 10,000 test cases, to evaluate the effectiveness of ChatTSS. Experimental results demonstrate that ChatTSS exhibits strong simplification performance across multiple evaluation dimensions, confirming its potential as an efficient and scalable TSR solution.
Gaolei Yi, Yuan Zhao 0010, Runkang Feng, Quanjun Zhang, Zhenyu Chen 0001
QRS1
2022 A Framework for Scanning Privacy Information based on Static Analysis
abstract
Modern software brings many conveniences to users through big data, but it also risks privacy leakage. In recent years, privacy leaks have been frequent, and various countries have introduced privacy protection bills to protect users' privacy security and avoid misuse of their private data.The researchers have conducted many studies to protect user privacy, including privacy policy compliance checks and mobile application permission checks. However, little existing work considers the verification of matching software code behavior and privacy policy. In this paper, we propose a set of privacy scanning methods to solve mentioned issues with static code analysis.We first classify privacy text and extracts privacy information. Then we perform static analysis on the code to obtain variable privacy information and privacy propagation paths by combining an abstract syntax tree and the call graph. We also match the results to the text analysis results. The experiments demonstrate that our method outperforms other classification methods in privacy text judgment, with an accuracy rate of 90% in detecting privacy information in the code. Meanwhile, the short running time ensures that no extra overhead is imposed on the user.
Yuan Zhao 0010, Gaolei Yi, Zhanwei Hui
QRS2