VLDB 2026 Research / reviewers in the wild / expert
Xiuting Ge
dblp:251/1990
· DBLP profile ↗
12ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0003-3683-7374ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Static code analyzer recommendation via preference mining
Xiuting Ge, Chunrong Fang, Xuanye Li, Ye Shang, Ya Pan |
Expert Syst. Appl. | 1 |
| 2025 | A Large-Scale Empirical Study of Actionable Warning Distribution Within ProjectsabstractStatic Analysis Tools (SATs) show potential defect detection ability while their usability is severely hindered by massive unactionable warnings. To improve the usability of SATs, many machine learning-based Actionable Warning Identification (AWI) studies have been proposed, which mainly focus on mining warning features and improving identification models to identify actionable warnings. However, the underlying distribution of the warning dataset, which is closely related to feature mining and thereby affects AWI model performance, is not well-explored by these studies. Further, there is a lack of a well-prepared warning dataset to support the distribution analysis. In this article, we first propose a warning dataset construction approach, which incorporates manual inspection and verification latency into postprocess labels from an advanced closed-warning heuristic and thereby acquire credible labels. Based on 10 large-scale and real-world projects with 25K+ revisions and 2087K+ SpotBugs warnings, we construct a qualified warning dataset with 11975 distinct warnings. Subsequently, we thoroughly analyze the actionable warning distribution within projects against our constructed dataset from six warning characteristics (i.e., category, type, priority, rank, file, and method). Based on the analysis results, we present 16 findings. Finally, a preliminary study demonstrates that our findings can be practical and instructive in improving the usability of SATs. Xiuting Ge, Chunrong Fang, Xuanye Li, Jia Liu 0015, Zhenyu Chen 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | CooTest: An Automated Testing Approach for V2X Communication SystemsabstractPerceiving the complex driving environment precisely is crucial to the safe operation of autonomous vehicles. With the tremendous advancement of deep learning and communication technology, Vehicle-to-Everything (V2X) collaboration has the potential to address limitations in sensing distant objects and occlusion for a single-agent perception system. However, despite spectacular progress, several communication challenges can undermine the effectiveness of multi-vehicle cooperative perception. The low interpretability of Deep Neural Networks (DNNs) and the high complexity of communication mechanisms make conventional testing techniques inapplicable for the cooperative perception of autonomous driving systems (ADS). Besides, the existing testing techniques, depending on manual data collection and labeling, become time-consuming and prohibitively expensive. In this paper, we design and implement CooTest, the first automated testing tool of the V2X-oriented cooperative perception module. CooTest devises the V2X-specific metamorphic relation and equips communication and weather transformation operators that can reflect the impact of the various cooperative driving factors to produce transformed scenes. Furthermore, we adopt a V2X-oriented guidance strategy for the transformed scene generation process and improve testing efficiency. We experiment CooTest with multiple cooperative perception models with different fusion schemes to evaluate its performance on different tasks. The experiment results show that CooTest can effectively detect erroneous behaviors under various V2X-oriented driving conditions. Also, the results confirm that CooTest can improve detection average precision and decrease misleading cooperation errors by retraining with the generated scenes. An Guo 0002, Zhenyu Chen 0001, Yuan Xiao 0003, Jiakai Liu, Xiuting Ge, Weisong Sun, Chunrong Fang |
ISSTA | 6 |
| 2024 | Improving actionable warning identification via the refined warning-inducing context representation
Xiuting Ge, Chunrong Fang, Xuanye Li, Quanjun Zhang, Jia Liu 0015, Zhenyu Chen 0001 |
Sci. China Inf. Sci. | 1 |
| 2024 | Pattern Mining-Based Warning Prioritization by Refining Abstract Syntax TreeabstractStatic code analysis tools (SATs) are widely used to detect potential defects in software projects. However, the usability of SATs is seriously hindered by a large number of unactionable warnings. Currently, many warning prioritization approaches are proposed to improve the usability of SATs. These approaches mainly extract different warning features to capture the statistical or historical information of warnings, thereby ranking actionable warnings in front of unactionable warnings. Such features are extracted by extremely relying on domain knowledge. However, the precise domain knowledge is difficult to be acquired. Also, the domain knowledge obtained in a project cannot be directly applied to other projects due to different application scenarios among different projects. To address the above problem, we propose a pattern mining-based warning prioritization approach based on the warning-related Abstract Syntax Tree (AST). To automatically mine actionable warning patterns, our approach leverages an advanced technique to collect actionable warnings, designs an algorithm to extract the warning-related AST, and mines patterns from ASTs of all actionable warnings. To prioritize the newly reported warnings, our approach combines exact and fuzzing matching techniques to calculate the similarity score between patterns of the newly reported warnings and the mined actionable warning patterns. We compare our approach with four typical baselines on five open-source and large-scale Java projects. The results show that our approach outperforms four baselines and achieves the maximum MAP (0.76) and MRR (2.19). Besides, a case study on Defect4J dataset demonstrates that our approach can discover 83% of true defects in the top 10 warnings. Xiuting Ge, Xuanye Li, Mingshuang Qing, Huibin Zhang, Xianyu Wu |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | A Survey of Source Code Search: A 3-Dimensional Perspectiveabstract(Source) code search is widely concerned by software engineering researchers because it can improve the productivity and quality of software development. Given a functionality requirement usually described in a natural language sentence, a code search system can retrieve code snippets that satisfy the requirement from a large-scale code corpus, e.g., GitHub. To realize effective and efficient code search, many techniques have been proposed successively. These techniques improve code search performance mainly by optimizing three core components, including query understanding component, code understanding component, and query-code matching component. In this article, we provide a 3-dimensional perspective survey for code search. Specifically, we categorize existing code search studies into query-end optimization techniques, code-end optimization techniques, and match-end optimization techniques according to the specific components they optimize. These optimization techniques are proposed to enhance the performance of specific components, and thus the overall performance of code search. Considering that each end can be optimized independently and contributes to the code search performance, we treat each end as a dimension. Therefore, this survey is 3-dimensional in nature, and it provides a comprehensive summary of each dimension in detail. To understand the research trends of the three dimensions in existing code search studies, we systematically review 68 relevant literatures. Different from existing code search surveys that only focus on the query end or code end or introduce various aspects shallowly (including codebase, evaluation metrics, modeling technique, etc.), our survey provides a more nuanced analysis and review of the evolution and development of the underlying techniques used in the three ends. Based on a systematic review and summary of existing work, we outline several open challenges and opportunities at the three ends that remain to be addressed in future work. Weisong Sun, Chunrong Fang, Yifei Ge, Yuling Hu, Quanjun Zhang, Xiuting Ge, Yang Liu 0003, Zhenyu Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2023 | An unsupervised feature selection approach for actionable warning identification
Xiuting Ge, Chunrong Fang, Jia Liu 0015, Mingshuang Qing, Xuanye Li |
Expert Syst. Appl. | 1 |
| 2023 | Test case classification via few-shot learning
Yuan Zhao 0010, Sining Liu, Quanjun Zhang, Xiuting Ge, Jia Liu 0008 |
Inf. Softw. Technol. | 4 |
| 2023 | An Empirical Study of Class Rebalancing Methods for Actionable Warning IdentificationabstractActionable warning identification (AWI) is crucial for improving the usability of static analysis tools. Currently, machine learning (ML)-based AWI approaches are notably common, which mainly focus on seeking high performance by improving the warning feature extraction and advancing the AWI model training. However, these approaches ignore an important fact that the number of actionable warnings is much smaller than that of unactionable warnings in the warning dataset used for the AWI model training (i.e., the class imbalance). Learning from such an imbalanced dataset may limit the performance of ML-based AWI approaches. To bridge the above gap, we are the first to conduct a comprehensive empirical study to investigate the impact of class imbalance on the ML-based AWI performance, whether class rebalancing methods can improve the ML-based AWI performance, and the differences of class rebalancing methods in the ML-based AWI model. Our empirical study is performed on 9 real-world and large-scale warning datasets, 25 typical class rebalancing methods, and 7 commonly used ML models. The experimental results show that 1) the class imbalance has a negative impact on the ML-based AWI performance; 2) 85% class rebalancing methods can significantly improve the ML-based AWI performance, but 8% ones do not work in the imbalanced warning datasets; 3) RandomOverSampler combined with AdaBoost/Random Forest can make the ML-based AWI model achieve optimal performance on nine warning datasets. Finally, we provide three practical guidelines that could help refine ML-based AWI approaches. Xiuting Ge, Chunrong Fang, Tongtong Bai, Jia Liu 0015 |
IEEE Trans. Reliab. | 1 |
| 2023 | Leveraging Android Automated Testing to Assist Crowdsourced TestingabstractCrowdsourced testing is an emerging trend in mobile application testing. The openness of crowdsourced testing provides a promising way to conduct large-scale and user-oriented testing scenarios on various mobile devices, while it also brings a problem, i.e., crowdworkers with different levels of testing experience severely threaten the quality of crowdsourced testing. Currently, many approaches have been proposed and studied to improve crowdsourced testing. However, these approaches do not fundamentally improve the ability of crowdworkers. In essence, the low-quality crowdsourced testing is caused by crowdworkers who are unfamiliar with the App Under Test (AUT) and do not know which part of the AUT should be tested. To address this problem, we propose a testing assistance approach, which leverages Android automated testing (i.e., dynamic and static analysis) to improve crowdsourced testing. Our approach constructs an Annotated Window Transition Graph (AWTG) model for the AUT by merging dynamic and static analysis results. Based on the AWTG model, our approach implements a testing assistance pipeline that provides the test task extraction, test task recommendation, and test task guidance to assist crowdworkers in testing the AUT. We experimentally evaluate our approach on real-world AUTs. The quantitative results demonstrate that our approach can effectively and efficiently assist crowdsourced testing. Besides, the qualitative results from a user study confirm the usefulness of our approach. Xiuting Ge, Shengcheng Yu, Chunrong Fang |
IEEE Trans. Software Eng. | 1 |
| 2022 | Locality-based security bug report identification via active learning
Xiuting Ge, Chunrong Fang, Meiyuan Qian, Mingshuang Qing |
Inf. Softw. Technol. | 1 |
| 2021 | Impact of datasets on machine learning based methods in Android malware detection: an empirical studyabstractFor Android malware detection, machine learning-based (ML-based) methods show promising performance. However, limited studies are performed to investigate the impact of factors related to datasets on ML-based methods, while the performance of ML-based methods dramatically relies on datasets. To partially bridge the gap, we conduct an empirical study to investigate the impact of factors related to datasets on ML-based Android malware detection methods. By investigating dataset differences between real-world scenarios and experimental settings, we summarize three dataset factors (i.e., class imbalance, quality, and timelines) and assess the impact of these factors on ML-based Android malware detection methods. We conduct experiments on more than 11K benign and 17K malicious applications. The results show that these three dataset factors yield significant biases in the existing ML-based Android malware detection methods. Based on these results, we learn some lessons about assessing ML-based Android malware detection methods when taking dataset factors into account. Xiuting Ge, Zhanwei Hui |
QRS | 1 |