Yuan Zhao 0010

dblp:65/2105-10 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-1980-6277ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 An Empirical Study on Machine Learning-Based Risk Prediction for Petroleum Pipelines
abstract
Effective risk prediction is essential for ensuring the safe operation of petroleum pipelines within the framework of pipeline integrity management. This study empirically investigates the application of machine learning techniques to pipeline risk prediction. By performing a correlation analysis on a risk-related dataset, the study identifies key relationships between various features and risk events, laying the groundwork for model development. Five machine learning algorithms-Decision Tree, Random Forest, Support Vector Machine, Neural Network, and Gradient Boosting Decision Tree (GBDT)-are implemented and evaluated. Experimental results indicate that the GBDT model outperforms the others, achieving an accuracy of 91%, precision of 94%, and recall of 96%. After parameter tuning, the GBDT model achieves a significantly improved accuracy of 99%. These results demonstrate that the GBDT-based model can effectively classify pipeline risk levels into high, medium-high, medium, and low categories, offering a robust and data-driven basis for risk identification and management in petroleum pipeline systems.
Haikang Gao, Ye Shang, Gaolei Yi, Yuan Zhao 0010, Zhenyu Chen 0001
QRS4
2025 Chattss: Improving Test Suite Simplification Via Large Language Models
abstract
As a critical component of software testing activities, regression testing plays an indispensable role in ensuring the correctness of software systems after changes. With the increasing scale and complexity of modern software, a pressing challenge arises: how to efficiently select the most effective test cases from existing test suites for regression testing, thereby reducing the associated cost. Although numerous methods have been proposed for test suite reduction, most of them rely on the assumption that test cases are independent of each other. In this paper, we present ChatTSS, a novel test case simplification approach powered by LLM. Unlike conventional test suite reduction that only shrinks the size of the test suite without altering individual test cases, ChatTSSleverages the program analysis capabilities of LLM to decompose test cases into fine-grained test atoms. It then applies appropriate reduction algorithms to perform more precise and effective test suite simplification. We conducted experiments on seven open-source projects, comprising over 10,000 test cases, to evaluate the effectiveness of ChatTSS. Experimental results demonstrate that ChatTSS exhibits strong simplification performance across multiple evaluation dimensions, confirming its potential as an efficient and scalable TSR solution.
Gaolei Yi, Yuan Zhao 0010, Runkang Feng, Quanjun Zhang, Zhenyu Chen 0001
QRS2
2025 Improving Deep Assertion Generation via Fine-Tuning Retrieval-Augmented Pre-Trained Language Models
abstract
Unit testing validates the correctness of the units of the software system under test and serves as the cornerstone in improving software quality and reliability. To reduce manual efforts in writing unit tests, some techniques have been proposed to generate test assertions automatically, including Deep Learning (DL)-based, retrieval-based, and integration-based ones. Among them, recent integration-based approaches inherit from both DL-based and retrieval-based approaches and are considered state-of-the-art. Despite being promising, such integration-based approaches suffer from inherent limitations, such as retrieving assertions with lexical matching while ignoring meaningful code semantics and generating assertions with a limited training corpus. In this article, we propose a novel Retrieval-Augmented Deep Assertion Generation (RetriGen) approach based on a hybrid assertion retriever and a Pre-Trained Language Model (PLM)-based assertion generator. Given a focal-test, RetriGen first builds a hybrid assertion retriever to search for the most relevant test–assert pair from external codebases. The retrieval process takes both lexical similarity and semantical similarity into account via a token-based and an embedding-based retriever, respectively. RetriGen then treats assertion generation as a sequence-to-sequence task and designs a PLM-based assertion generator to predict a correct assertion with historical test–assert pairs and the retrieved external assertion. Although our concept is general and can be adapted to various off-the-shelf encoder–decoder PLMs, we implement RetriGen to facilitate assertion generation based on the recent CodeT5 model. We conduct extensive experiments to evaluate RetriGen against six state-of-the-art approaches across two large-scale datasets and two metrics. The experimental results demonstrate that RetriGen achieves 57.66% and 73.24% in terms of accuracy and CodeBLEU, outperforming all baselines with an average improvement of 50.66% and 14.14%, respectively. Furthermore, RetriGen generates 1,598 and 1,818 unique correct assertions that all baselines fail to produce, 3.71X and 4.58X more than the most recent approach EditAS . We also demonstrate that adopting other PLMs can provide substantial advancement, e.g., four additionally utilized PLMs outperform EditAS by 7.91%–12.70% accuracy improvement, indicating the generalizability of RetriGen. Overall, our study highlights the promising future of fine-tuning off-the-shelf PLMs to generate accurate assertions by incorporating external knowledge sources.
Quanjun Zhang, Chunrong Fang, Yuan Zhao 0010, Rubing Huang, Yun Yang 0001, Tao Zheng 0005, Zhenyu Chen 0001
ACM Trans. Softw. Eng. Methodol.5
2025 Improving Retrieval-Augmented Deep Assertion Generation via Joint Training
abstract
Unit testing attempts to validate the correctness of basic units of the software system under test and has a crucial role in software development and testing. However, testing experts have to spend a huge amount of effort to write unit test cases manually. Very recent work proposes a retrieve-and-edit approach to automatically generate unit test oracles,i.e.,assertions. Despite being promising, it is still far from perfect due to some limitations, such as splitting assertion retrieval and generation into two separate components without benefiting each other. In this paper, we propose AG-RAG, a retrieval-augmented automated assertion generation (AG) approach that leverages external codebases and joint training to address various technical limitations of prior work. Inspired by the plastic surgery hypothesis, AG-RAG attempts to combine relevant unit tests and advanced pre-trained language models (PLMs) with retrieval-augmented fine-tuning. The key insight of AG-RAG is to simultaneously optimize the retriever and the generator as a whole pipeline with a joint training strategy, enabling them to learn from each other. Particularly, AG-RAG builds a dense retriever to search for relevant test-assert pairs (TAPs) with semantic matching and a retrieval-augmented generator to synthesize accurate assertions with the focal-test and retrieved TAPs as input. Besides, AG-RAG leverages a code-aware language model CodeT5 as the cornerstone to facilitate both assertion retrieval and generation tasks. Furthermore, AG-RAG designs a joint training strategy that allows the retriever to learn from the feedback provided by the generator. This unified design fully adapts both components specifically for retrieving more useful TAPs, thereby generating accurate assertions. AG-RAG is a generic framework that can be adapted to various off-the-shelf PLMs. We extensively evaluate AG-RAG against six state-of-the-art AG approaches on two benchmarks and three metrics. Experimental results show that AG-RAG significantly outperforms previous AG approaches on all benchmarks and metrics,e.g.,improving the most recent baselineEditASby 20.82% and 26.98% in terms of accuracy. AG-RAG also correctly generates 1739 and 2866 unique assertions that all baselines fail to generate, 3.45X and 9.20X more thanEditAS. We further demonstrate the positive contribution of our joint training strategy,e.g.,AG-RAG improving a variant without the retriever by an average accuracy of 14.11%. Besides, adopting other PLMs can provide substantial advancement,e.g.,AG-RAG with four different PLMs improving EditAS by an average accuracy of 9.02%, highlighting the generalizability of our framework. Overall, our work demonstrates the promising potential of jointly fine-tuning the PLM-based retriever and generator to predict accurate assertions by incorporating external knowledge sources, thereby reducing the manual efforts of unit testing experts in practical scenarios.
Quanjun Zhang, Chunrong Fang, Ruixiang Qian, Shengcheng Yu, Yuan Zhao 0010, Yun Yang 0001, Tao Zheng 0005, Zhenyu Chen 0001
IEEE Trans. Software Eng.6
2023 Test case classification via few-shot learning
Yuan Zhao 0010, Sining Liu, Quanjun Zhang, Xiuting Ge, Jia Liu 0008
Inf. Softw. Technol.1
2022 A Framework for Scanning Privacy Information based on Static Analysis
abstract
Modern software brings many conveniences to users through big data, but it also risks privacy leakage. In recent years, privacy leaks have been frequent, and various countries have introduced privacy protection bills to protect users' privacy security and avoid misuse of their private data.The researchers have conducted many studies to protect user privacy, including privacy policy compliance checks and mobile application permission checks. However, little existing work considers the verification of matching software code behavior and privacy policy. In this paper, we propose a set of privacy scanning methods to solve mentioned issues with static code analysis.We first classify privacy text and extracts privacy information. Then we perform static analysis on the code to obtain variable privacy information and privacy propagation paths by combining an abstract syntax tree and the call graph. We also match the results to the text analysis results. The experiments demonstrate that our method outperforms other classification methods in privacy text judgment, with an accuracy rate of 90% in detecting privacy information in the code. Meanwhile, the short running time ensures that no extra overhead is imposed on the user.
Yuan Zhao 0010, Gaolei Yi, Zhanwei Hui
QRS1
2020 Test recommendation system based on slicing coverage filtering
abstract
Software testing plays a crucial role in software lifecycle. As a basic approach of software testing, unit testing is one of the necessary skills for software practitioners. Since testers are required to understand the inner code of the software under test(SUT) while writing a test case, testers usually need to learn how to detect the bug within SUT effectively. When novice programmers started to learn writing unit tests, they will generally watch a video lesson or reading unit tests written by others. These learning approaches are either time-consuming or too hard for a novice. To solve these problems, we developed a system, named TeSRS, to assist novice programmers to learn unit testing. TeSRS is a test recommendation system which can effectively assist test novice in learning unit testing. Utilizing program slice technique, TeSRS has gotten an enormous amount of test snippets from superior crowdsourcing test scripts. Depending on these test snippets, TeSRS provides novices a easier way for unit test learning. To sum up, TeSRS can help test novices (1) obtain high level design ideas of unit test case and (2) improve capabilities(e.g. branch coverage rate and mutation coverage rate) of their test scripts. TeSRS has built a scalable corpus composed of over 8000 test snippets from more than 25 test problems. Its stable performance shows effectiveness in unit test learning.
Ruixiang Qian, Yuan Zhao 0010, Duo Men, Yang Feng 0003, Qingkai Shi, Zhenyu Chen 0001
ISSTA2
2020 Quality assessment of crowdsourced test cases
Yuan Zhao 0010, Yang Feng 0003, Yi Wang 0013, Chunrong Fang, Zhenyu Chen 0001
Sci. China Inf. Sci.1
2019 Towards Generating Cost-Effective Test-Suite for Ethereum Smart Contract
abstract
In Ethereum, many accounts and funds have been managed by smart contracts, thereby making them easy to be targeted. Due to the persistence characteristic of blockchain, revising a deployed smart contract is almost impossible. Both realities heighten the risks of managing funds and thus increase the demand for conducting sufficient testing to Ethereum Smart Contracts (ESC). Different from the conventional software, ESC is a gas-driven program, where developers must charge gases for deploying and testing it. Therefore, it is important to provide a cost-effective yet representative test suite, where its representativeness can be typically measured by its branch coverage. In this paper, we deem the problem of ESC test generation as a Pareto minimization problem, and three objectives, minimizing (1) uncovered branch coverage, (2) time cost, and (3) gas cost are considered. Then, we propose a random based and an NSGA-II based multi-objective approach to seek cost-effective test-suites. Our empirical study on a set of smart contracts in eight of the most widely used Ethereum Decentralized Applications (DApps) verified that the proposed approaches could significantly reduce the gas cost as well as the time cost while retaining the ability to cover branches.
Xingya Wang, Weisong Sun, Yuan Zhao 0010
SANER4
2019 A Unified Framework for Bug Report Assignment
abstract
It is typically a manual, time-consuming, and tedious task of assigning bug reports to individual developers. Although some machine learning techniques are adopted to alleviate this dilemma, they are mainly focused on the open source projects, which use traditional repositories such as Bugzilla to manage their bug reports. With the boom of the mobile Internet, some new requirements and methods of software testing are emerging, especially the crowdsourced testing. Unlike the traditional channels, whose bug reports are often heavyweight, which means their bug reports are standardized with detailed attribute localization, bug reports tend to be lightweight in the context of crowdsourced testing. To exploit the differences of the bug reports assignment in the new settings, a unified bug reports assignment framework is proposed in this paper. This framework is capable of handling both the traditional heavyweight bug reports and the lightweight ones by (i) first preprocessing the bug reports and feature selections, (ii) then tuning the parameters that indicate the ratios of choosing different methods to vectorize bug reports, (iii) and finally applying classification algorithms to assign bug reports. Extensive experiments are conducted on three datasets to evaluate the proposed framework. The results indicate the applicability of the proposed framework, and also reveal the differences of bug report assignment between traditional repositories and crowdsourced ones.
Yuan Zhao 0010, Tieke He, Zhenyu Chen 0001
Int. J. Softw. Eng. Knowl. Eng.1