Shan Gao 0009

dblp:67/4510-9 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0006-2695-2968ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 8 since 2021
YearPublicationVenuePosition
2026 Understanding the Potentially Confounding Effect of Test Suite Size in Test Effectiveness Evaluation
abstract
Background . Code coverage and mutation score serve as pivotal test effectiveness metrics used to assess a test suite’s ability to uncover actual defects. However, prior research has produced inconsistent or even conflicting findings regarding their correlation with defect detection capability, particularly concerning the impact of test suite size. Problem. The extent of the potentially confounding effect of test suite size in test effectiveness evaluation context is not clear, nor is the method to remove the potentially confounding effect, or the influence of this removal on the performance of test suite optimization. Objective . Our goal is to deeply understand how test suite size affects the true relationship between test effectiveness metrics and a test suite’s ability to detect actual defects. Method. We first employ statistical methods to examine the extent of the potentially confounding effect of test suite size in the context of test effectiveness evaluation. After that, we propose a linear regression-based method to remove the potentially confounding effect of test suite size. Finally, we empirically explore the impact of this removal method on test suite optimization. Result. Our experimental results, based on the Defects4J defect dataset, uncovers that: (1) the confounding effect of test suite size on the associations between test effectiveness metrics and defect detection capability in general exists; (2) the proposed linear regression-based method can effectively remove the confounding effect; and (3) after removing the confounding effect, mutation score demonstrates superior effectiveness in predicting test suite effectiveness, while statement coverage is the least effective metric. Furthermore, both coverage-based and mutation-based test suite reduction exhibit enhanced cost-effectiveness in defect detection, and there is a marginal improvement in the speed of defect detection for coverage-based test case prioritization. Conclusion . When using test effectiveness metrics to assess test suite effectiveness, it is crucial to eliminate the influence of test suite size.
Yang Wang 0165, Peng Zhang 0083, Shan Gao 0009, Yibiao Yang, Yanhui Li 0001, Lin Chen 0015, Yuming Zhou
ACM Trans. Softw. Eng. Methodol.7
2025 Multi-view Leaderboard: Towards Evaluating the Code Intelligence of LLMs From Multiple Views
abstract
Large Language Models (LLMs) have shown remarkable performance in code intelligence tasks, prompting the development of various benchmarks and leaderboards to assess their effectiveness across diverse programming scenarios. However, existing leaderboards often rely on coarse-grained metrics and overlook performance variations across different types of tasks. In this paper, we introduce Multi-view Leaderboard, a comprehensive evaluation framework designed to assess the coding capabilities of LLMs from multiple views. Our leaderboard partitions widely-used datasets such as HumanEval, MBPP, and ComplexCodeEval into subsets based on factors like prompt length, problem complexity, and task type. It supports four popular code intelligence tasks including code generation, code completion, test case generation, and API recommendation. Additionally, our leaderboard presents results using ranking tables, line charts, radar charts, and heatmaps. Based on LLMs’ performance on different subsets, we provide model recommendations tailored to different real-world scenarios via a Sankey diagram. A user study involving 11 participants revealed that 90% valued the leaderboard’s practical usefulness for analyzing LLMs’ code intelligence from multiple perspectives. The Multi-view Leaderboard is available at https://huggingface.co/spaces/MVLLL/Multi-view-leaderboard. The demonstration video is available at https://youtu.be/J-zQiOYa1Y8
Zexun Zhan, Cuiyun Gao 0001, Yujia Chen 0004, Guoai Xu, Chun Yong Chong, Shan Gao 0009, Xin Xia 0001
APSEC7
2025 Correction to: A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli
Empir. Softw. Eng.7
2025 The Current Challenges of Software Engineering in the Era of Large Language Models
abstract
With the advent of large language models (LLMs) in the AI area, the field of software engineering (SE) has also witnessed a paradigm shift. These models, by leveraging the power of deep learning and massive amounts of data, have demonstrated an unprecedented capacity to understand, generate, and operate programming languages. They can assist developers in completing a broad spectrum of software development activities, encompassing software design, automated programming, and maintenance, which potentially reduces huge human efforts. Integrating LLMs within the SE landscape (LLM4SE) has become a burgeoning trend, necessitating exploring this emergent landscape’s challenges and opportunities. The article aims at revisiting the software development lifecycle (SDLC) under LLMs, and highlighting challenges and opportunities of the new paradigm. The article first summarizes the overall process of LLM4SE, and then elaborates on the current challenges based on a through discussion. The discussion was held among more than 20 participants from academia and industry, specializing in fields such as SE and artificial intelligence. Specifically, we achieve 26 key challenges from seven aspects, including software requirement and design, coding assistance, testing code generation, code review, code maintenance, software vulnerability management, and data, training, and evaluation. We hope the achieved challenges would benefit future research in the LLM4SE field.
Cuiyun Gao 0001, Xing Hu 0008, Shan Gao 0009, Xin Xia 0001, Zhi Jin 0001
ACM Trans. Softw. Eng. Methodol.3
2024 Mining Pull Requests to Detect Process Anomalies in Open Source Software Development
abstract
Trustworthy Open Source Software (OSS) development processes are the basis that secures the long-term trustworthiness of software projects and products. With the aim to investigate the trustworthiness of the Pull Request (PR) process, the common model of collaborative development in OSS community, we exploit process mining to identify and analyze the normal and anomalous patterns of PR processes, and propose our approach to identifying anomalies from both control-flow and semantic aspects, and then to analyze and synthesize the root causes of the identified anomalies. We analyze 17531 PRs of 18 OSS projects on GitHub, extracting 26 root causes of control-flow anomalies and 19 root causes of semantic anomalies. We find that most PRs can hardly contain both semantic anomalies and control-flow anomalies, and the internal custom rules in projects may be the key causes for the identified anomalous PRs. We further discover and analyze the patterns of normal PR processes. We find that PRs in the non-fork model (42%) are far more likely than the fork model (5%) to bypass the review process, indicating a higher potential risk. Besides, we analyzed nine poisoned projects whose PR practices were indeed worse. Given the complex and diverse PR processes in OSS community, the proposed approach can help identify and understand not only anomalous PRs but also normal PRs, which offers early risk indications of suspicious incidents (such as poisoning) to OSS supply chain.
Bohan Liu 0003, He Zhang 0001, Weigang Ma, Hongyu Kuang, Jinwei Xu, Shan Gao 0009
ICSE7
2024 ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
abstract
In recent years, with the widespread attention of academia and industry on the application of large language models (LLMs) to code-related tasks, an increasing number of large code models (LCMs) have been proposed and corresponding evaluation benchmarks have continually emerged. Although existing evaluation benchmarks are helpful for comparing different LCMs, they may not reflect the performance of LCMs in various development scenarios. Specifically, they might evaluate model performance in only one type of scenario (e.g., code generation or code completion), whereas real development contexts are diverse and may involve multiple tasks such as code generation, code completion, API recommendation, and test function generation. Additionally, the questions may not originate from actual development practices, failing to capture the programming challenges faced by developers during the development process.
Cuiyun Gao 0001, Chun Yong Chong, Chaozheng Wang, Shan Gao 0009, Xin Xia 0001
ASE6
2024 A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
abstract
API suggestion is a critical task in modern software development, assisting programmers by predicting and recommending third-party APIs based on the current context. Recent advancements in large code models (LCMs) have shown promise in the API suggestion task. However, they mainly focus on suggesting which APIs to use, ignoring that programmers may demand more assistance while using APIs in practice including when to use the suggested APIs and how to use the APIs. To mitigate the gap, we conduct a systematic evaluation of LCMs for the API suggestion task in the paper.
Chaozheng Wang, Shuzheng Gao, Cuiyun Gao 0001, Wenxuan Wang 0001, Chun Yong Chong, Shan Gao 0009, Michael R. Lyu
ASE6
2024 A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli
Empir. Softw. Eng.7