VLDB 2026 Research / reviewers in the wild / expert
Xiaoxue Ren
dblp:187/9569
· DBLP profile ↗
28ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 10 first-author · 16 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution ReasoningabstractLingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye, Zhongxin Liu, Xiaoxue Ren, Lingfeng Bao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye, Zhongxin Liu 0002, Xiaoxue Ren, Lingfeng Bao |
ACL (1) | 6 |
| 2026 | The Bidirectional Process Reward ModelabstractProcess Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promising approach to enhance the reasoning quality of Large Language Models (LLMs).However, most existing PRMs rely on a unidirectional left-to-right (L2R) evaluation scheme, which restricts their utilization of global context.In light of this challenge, we propose a novel bidirectional evaluation paradigm, named Bidirectional Process Reward Model (BiPRM).BiPRM incorporates a parallel right-to-left (R2L) evaluation stream, implemented via prompt reversal, alongside the conventional L2R flow.Then a gating mechanism is introduced to adaptively fuse the reward scores from both streams to yield a holistic quality assessment.Remarkably, compared to the original PRM, BiPRM introduces only a 0.3% parameter increase for the gating module, and the parallel execution of two streams incurs merely 5% inference time latency.Our extensive empirical evaluations spanning diverse benchmarks, LLM backbones, PRM objectives and sampling policies demonstrate that BiPRM consistently surpasses unidirectional baselines, achieving an average relative gain of 10.6% over 54 solution-level configurations and 37.7% in 12 step-level error detection scenarios.Generally, our results highlight the effectiveness, robustness and general applicability of BiPRM, offering a promising new direction for processbased reward modeling.1 Lingyin Zhang, Xiaoxue Ren, Ziqiang Cao |
ACL (1) | 3 |
| 2026 | HMF: Enhancing reentrancy vulnerability detection and repair with a hybrid model frameworkabstractSmart contracts have revolutionized the credit landscape. However, their security remains intensely scrutinized due to numerous hacking incidents and inherent logical challenges. One well-known issue is reentrancy vulnerability, exemplified by DAO attacks that lead to substantial economic losses. Previous approaches have employed rule-based and deep learning-based (DL) algorithms to detect and repair reentrancy vulnerability. Large language models (LLM) have been distinguished in recent years for their excellent understanding of text and code. However, less attention has been paid to LLM-based reentrancy vulnerability detection and repair, and direct prompt-based approaches often suffer from inefficiencies and high false positives. To overcome the above shortcomings, this paper proposes a hybrid model framework combining LLM with DL to enhance the detection and repair of reentrancy vulnerabilities. This unified framework comprises three crucial phases: the data processing phase, the vulnerability detection phase, and the vulnerability repair phase. Extensive experimental results validate the superiority of our approach over state-of-the-art baselines, and ablation studies demonstrate the effectiveness of each component. Our approach demonstrates significant improvements in vulnerability detection, with increases of 3.51% in accuracy, 2.31% in recall, 0.42% in precision, and 0.85% in F1-score. Furthermore, our approach can achieve a notable 9.62% enhancement in the repair rate. Finally, we also conducted a user study to emphasize its potential to fortify the security of smart contracts. Mengliang Li, Xiaoxue Ren, Zhuo Li 0014, Jianling Sun |
Autom. Softw. Eng. | 3 |
| 2026 | AdaCoder: An Adaptive Planning and Multi-Agent Framework for Function-Level Code GenerationabstractRecently, researchers have proposed many multi-agent frameworks for function-level code generation, which aim to improve software development productivity by automatically generating function-level source code based on task descriptions. A typical multi-agent framework consists of Large Language Model (LLM)-based agents that are responsible for task planning, code generation, testing, debugging, etc. Studies have shown that existing multi-agent code generation frameworks perform well on ChatGPT. However, their generalizability across other foundation LLMs remains unexplored systematically. In this paper, we report an empirical study on the generalizability of four state-of-the-art multi-agent code generation frameworks across 12 open-source LLMs with varying code generation and instruction-following capabilities. Our study reveals the unstable generalizability of existing frameworks on diverse foundation LLMs. Based on the findings obtained from the empirical study, we propose AdaCoder, a novel adaptive planning, multi-agent framework for function-level code generation. AdaCoder has two phases. Phase-1 is an initial code generation step without planning, which uses an LLM-based coding agent and a script-based testing agent to unleash LLM’s native power, identify cases beyond LLM’s power, and determine the errors hindering execution. Phase-2 adds a rule-based debugging agent and an LLM-based planning agent for iterative code generation with planning. Our evaluation shows that AdaCoder achieves higher generalizability on diverse LLMs. Compared to the best baseline MapCoder, AdaCoder is on average 27.69% higher in Pass@1, 16 times faster in inference, and 12 times lower in token consumption. Yueheng Zhu, Chao Liu 0014, Xiaoxue Ren, Zhongxin Liu 0002, Ruwei Pan, Hongyu Zhang 0002 |
IEEE Trans. Software Eng. | 4 |
| 2025 | Bridging Solidity Evolution Gaps: An LLM-Enhanced Approach for Smart Contract Compilation Error ResolutionabstractSolidity, the dominant smart contract language for Ethereum, has rapidly evolved with frequent version updates to enhance security, functionality, and developer experience. However, these continual changes introduce significant challenges, particularly in compilation errors, code migration, and maintenance. Therefore, we conduct an empirical study to investigate the challenges in the Solidity version evolution and reveal that 81.68 % of examined contracts encounter errors when compiled across different versions, with 86.92 % of compilation errors. To mitigate these challenges, we conducted a systematic evaluation of large language models (LLMs) for resolving Solidity compilation errors during version migrations. Our empirical analysis across both open-source (LLaMA3, DeepSeek) and closedsource (GPT-4o, GPT-3.5-turbo) LLMs reveals that although these models exhibit error repair capabilities, their effectiveness diminishes significantly for semantic-level issues and shows strong dependency on prompt engineering strategies. This underscores the critical need for domain-specific adaptation in developing reliable LLM-based repair systems for smart contracts. Building upon these insights, we introduce SMCFIXER, a novel framework that systematically integrates expert knowledge retrieval with LLM-based repair mechanisms for Solidity compilation error resolution. The architecture comprises three core phases: (1) context-aware code slicing that extracts relevant error information; (2) expert knowledge retrieval from official documentation; and (3) iterative patch generation for Solidity migration. Experimental validation across Solidity version migrations demonstrates our approach's statistically significant 24.24% improvement over baseline GPT-4o on real-world datasets, achieving near-perfect 96.97% accuracy. Likai Ye, Mengliang Li, Dehai Zhao, Jiamou Sun, Xiaoxue Ren |
ICSME | 5 |
| 2025 | Issue Localization via LLM-Driven Iterative Code Graph SearchingabstractIssue solving aims to generate patches to fix re-ported issues in real-world code repositories according to issue descriptions. Issue localization forms the basis for accurate issue solving. Recently, large language model (LLM) based issue localization methods have demonstrated state-of-the-art performance. However, these methods either search from files mentioned in issue descriptions or in the whole repository and struggle to balance the breadth and depth of the search space to converge on the target efficiently. Moreover, they allow LLM to explore whole repositories freely, making it challenging to control the search direction to prevent the LLM from searching for incorrect targets. Meanwhile, because LLMs may not correctly produce the required interaction formats with the environment, they suffer from search failures.This paper introduces COSIL, an LLM-driven, powerful function-level issue localization method without training or indexing. To balance search breadth and depth, COSIL employs a two-phase code graph search strategy. It first conducts broad exploration at the file level using dynamically constructed module call graphs, and then performs in-depth analysis at the function level by expanding the module call graph into a function call graph and executing iterative searches. To precisely control the search direction, COSIL designs a pruner to filter unrelated directions and irrelevant contexts. To avoid incorrect interaction formats in long contexts, COSIL introduces a reflection mechanism that uses additional independent queries in short contexts to enhance formatted abilities. Experiment results demonstrate that COSIL achieves a Top-1 localization accuracy of 43.3% and 44.6% on SWE-bench Lite and SWE-bench Verified, respectively, with Qwen2.5-Coder-32B, average outperforming the state-of-the-art methods by 96.04%. When COSIL is integrated into an issue-solving method, Agentless, the issue resolution rate improves by 2.98%–30.5%. Zhonghao Jiang, Xiaoxue Ren, Meng Yan 0001, Wei Jiang 0041, Yong Li 0004, Zhongxin Liu 0002 |
ASE | 2 |
| 2025 | PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code EditingabstractLarge Language Models (LLMs) have demonstrated significant capability in code generation, but their potential in code efficiency optimization remains underexplored. Previous LLM-based code efficiency optimization approaches exclusively focus on function-level optimization and overlook interaction between functions, failing to generalize to real-world development scenarios. Code editing techniques show great potential for conducting project-level optimization, yet they face challenges associated with invalid edits and suboptimal internal functions. To address these gaps, we propose PEACE, a novel hybrid framework for Project-Level code Efficiency optimization through Automatic Code Editing, which also ensures the overall correctness and integrity of the project. PEACE integrates three key phases: dependency-aware optimizing function sequence construction, valid associated edits identification, and efficiency optimization editing iteration. To rigorously evaluate the effectiveness of PEACE, we construct PEACEXEC, the first benchmark comprising 146 real-world optimization tasks from 47 high-impact GitHub Python projects, along with highly qualified test cases and executable environments. Extensive experiments demonstrate PEACE’s superiority over the state-of-the-art baselines, achieving a 69.2% correctness rate (pass@1), +46.9% opt rate, and 0.840 speedup in execution efficiency. Notably, our PEACE outperforms all baselines by significant margins, particularly in complex optimization tasks with multiple functions. Moreover, extensive experiments are also conducted to validate the contributions of each component in PEACE, as well as the rationale and effectiveness of our hybrid framework design. Xiaoxue Ren, Yun Peng 0003, Zhongxin Liu 0002, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
ASE | 1 |
| 2025 | A Hybrid Attention-Based Fuzzy Pooling Network Model for Locating Polyp Positions in Gastroscopic Image in Internet of Medical ThingsabstractUsing deep-learning techniques, such as convolutional neural networks to locate polyps in gastroscopic images automatically is of great significance for preventing gastric cancer. However, polyps are typically irregular in shape and inconsistent in size, posing challenges to the predictive performance of previous models. Previous localization models typically treat all parts of the gastroscopic image equally with multiple convolution and pooling operations without considering the importance of different positions. As polyps may occupy only a tiny portion of the gastroscopic image, the model should prioritize attention to different positions in the image differently. Moreover, in previous models, pooling operations generally employ max-pooling or average-pooling, which are less effective in information retention and positional awareness. Therefore, inspired by attention mechanisms and fuzzy algorithms, in this work, we propose a Hybrid Attention Fuzzy Pooling Network (HAFPN) for locating polyps in gastroscopic images in Internet of Medical Things. The advantages of the HAFPN model lie in the following. First, valuable local information in the image is preserved using a hybrid attention mechanism. Specifically, the HAFPN model inputs gastroscopic images into a convolutional neural network augmented with channel attention and spatial attention mechanisms to calculate the coordinates of gastric polyps. The combination of channel attention and spatial attention mechanisms helps the HAFPN model dynamically capture channel and spatial information from the original image. Second, fuzzy pooling operations preserve positional information in the feature maps. We conducted extensive experiments on a real gastroscopic dataset to validate the effectiveness of the HAFPN model. Honghu Wang, Xiaoxue Ren |
IEEE Internet Things J. | 3 |
| 2025 | KG4VA: Constructing Vulnerability Knowledge Graph for Software Vulnerability AssessmentabstractSoftware vulnerabilities pose serious threats to software security. When faced with multiple software vulnerabilities at the same time, it is urgent to determine whether the vulnerabilities are high-risk. Existing vulnerability assessment approaches only learn the mapping relationships between vulnerability descriptions and severity levels, while ignoring the sharing of the same or similar elements between vulnerabilities. Furthermore, solely focusing on vulnerability descriptions fails to accurately characterize the vulnerability behavior. In this paper, we propose a novel vulnerability knowledge graph (KG) to capture the relationships between vulnerabilities. To construct the vulnerability KG automatically, we propose to leverage vulnerability elements extracted from vulnerability descriptions to link different vulnerabilities. Based on the constructed KG, we further propose a novel KG-based vulnerability assessment (VA) approach KG4VA, which precisely finds the similar vulnerability for an encountered vulnerability description by analyzing and matching the elements entities based on the vulnerability KG. The experiment results show that KG4VA outperforms the baselines in almost all metrics (e.g., 3.27%-10.83% accuracy improvements). Moreover, our ablation experiments demonstrate that the vulnerability knowledge graph can indeed offer valuable information for vulnerability assessment. Zhenlei Ye, Xiaobing Sun 0001, Lili Bo, Sicong Cao, Xiaoxue Ren, Lianyong Qi, Jiale Zhang 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Hydra-Reviewer: A Holistic Multi-Agent System for Automatic Code Review Comment GenerationabstractReview comment generation is a crucial task in code review, and significant progress has been made in automating. Previous research has generated review comments by fine-tuning pre-trained models or Large Language Models (LLMs). However, these studies have overlooked the necessity of conducting code reviews from multiple perspectives, resulting in the omission of potential issues in code changes. Additionally, the complexity of review comments often hinders the accurate quantitative evaluation of automated tools’ effectiveness.In this paper, we first conduct an empirical study to propose a comprehensive taxonomy of code review dimensions. We also identify three major limitations of existing automated code review (ACR) methods: lack of comprehensiveness, incorrectness, and vagueness. Building on the insights from our empirical study, we introduce Hydra-Reviewer, a collaborative multi-agent framework powered by large language models, designed to automatically generate high-quality code reviews. We utilize the CodeReview and CodeReviewNew benchmark datasets, along with a newly constructed review comment generation dataset. We compare Hydra-Reviewerwith several baselines, including CodeReviewer, LLaMA-Reviewer, ChatGPT, Comprehensive-ChatGPT, and DeepSeek-V3.The experimental results show that Hydra-Reviewerachieves a BLEU score of 8.20, outperforming the state-of-the-art baseline, DeepSeek-V3, which scores 7.85. In qualitative evaluation, Hydra-Reviewer’s generated comments span an average of 7.8 review dimensions, addressing the limitations of existing ACR methods effectively. Additionally, Hydra-Reviewerdemonstrates strong generalization capabilities on unseen dataset. We further validate the contributions of each component of Hydra-Reviewerthrough an ablation study and confirm the helpfulness and readability of the generated comments via a User Study. Finally, a cost analysis reveals that Hydra-Reviewergenerates review comments at an average cost of 0.018 dollars and 62.63 seconds per code change. Xiaoxue Ren, Chaoqun Dai, Ye Wang 0012, Chao Liu 0014, Bo Jiang 0009 |
IEEE Trans. Software Eng. | 1 |
| 2025 | FlexFL: Flexible and Effective Fault Localization With Open-Source Large Language ModelsabstractFault localization (FL) targets identifying bug locations within a software system, which can enhance debugging efficiency and improve software quality. Due to the impressive code comprehension ability of Large Language Models (LLMs), a few studies have proposed to leverage LLMs to locate bugs, i.e., LLM-based FL, and demonstrated promising performance. However, first, these methods are limited in flexibility. They rely on bug-triggering test cases to perform FL and cannot make use of other available bug-related information, e.g., bug reports. Second, they are built upon proprietary LLMs, which are, although powerful, confronted with risks in data privacy. To address these limitations, we propose a novel LLM-based FL framework named FlexFL, which can flexibly leverage different types of bug-related information and effectively work with open-source LLMs. FlexFL is composed of two stages. In the first stage, FlexFL reduces the search space of buggy code using state-of-the-art FL techniques of different families and provides a candidate list of bug-related methods. In the second stage, FlexFL leverages LLMs to delve deeper to double-check the code snippets of methods suggested by the first stage and refine fault localization results. In each stage, FlexFL constructs agents based on open-source LLMs, which share the same pipeline that does not postulate any type of bug-related information and can interact with function calls without the out-of-the-box capability. Extensive experimental results on Defects4J demonstrate that FlexFL outperforms the baselines and can work with different open-source LLMs. Specifically, FlexFL with a lightweight open-source LLM Llama3-8B can locate 42 and 63 more bugs than two state-of-the-art LLM-based FL approaches AutoFL and AgentFL that both use GPT-3.5. In addition, FlexFL can localize 93 bugs that cannot be localized by non-LLM-based FL techniques at the top 1. Furthermore, to mitigate potential data contamination, we conduct experiments on a dataset which Llama3-8B has not seen before, and the evaluation results show that FlexFL can also achieve good performance. Chuyang Xu, Zhongxin Liu 0002, Xiaoxue Ren, Gehao Zhang, David Lo 0001 |
IEEE Trans. Software Eng. | 3 |
| 2024 | JumpCoder: Go Beyond Autoregressive Coder via Online ModificationabstractWhile existing code large language models (code LLMs) exhibit impressive capabilities in code generation, their autoregressive sequential generation inherently lacks reversibility.This limitation hinders them from timely correcting previous missing statements during coding as humans do, often leading to error propagation and suboptimal performance.We introduce JUMPCODER, a novel model-agnostic framework that enables human-like online modification and non-sequential generation to augment code LLMs.The key idea behind JUMP-CODER is to insert new code into the currently generated code when necessary during generation, which is achieved through an auxiliary infilling model that works in tandem with the code LLM.Since identifying the best infill position beforehand is intractable, we adopt an infill-first, judge-later strategy, which experiments with filling at the k most critical positions following the generation of each line, and uses an Abstract Syntax Tree (AST) parser alongside the Generation Model Scoring to effectively judge the validity of each potential infill.Extensive experiments using six state-ofthe-art code LLMs across multiple and multilingual benchmarks consistently indicate significant improvements over all baselines.Our code is public at https://github.com/ Keytoyze/JumpCoder. Mouxiang Chen, Zhongxin Liu 0002, Xiaoxue Ren, Jianling Sun |
ACL (1) | 4 |
| 2024 | Enhancing Reentrancy Vulnerability Detection and Repair with a Hybrid Model Framework
Mengliang Li, Xiaoxue Ren, Zhuo Li 0014, Jianling Sun |
APSEC | 2 |
| 2024 | Dynamic NFT Classification and Detection on Ethereum via Smart ContractabstractIn recent years, Non-Fungible Token (NFT) has gradually become the key application of blockchain technology. Static NFT is the most common type of NFT. Once static NFT is minted on the blockchain, its additional metadata is immutable. However, some NFTs that mark real assets, games, sports, and other types must dynamically update the metadata. Therefore, a dynamic NFT with changeable features is needed. The emergence of dynamic NFT has greatly expanded the application innovation scene, and promoted the rapid development of community ecology, but also brought new problems and challenges to anti-fraud and supervision. This paper aims to realize the classification and detection of dynamic NFT. First, define and classify dynamic NFTs from both dynamic and static perspectives. Second, a complete dataset of dynamic NFT smart contract codes on Ethereum was constructed for the first time, and analyzed from multiple perspectives. Third, a smart contract feature model of dynamic NFT is proposed, and machine learning methods are used for recognition and classification. After experimental verification, the method proposed in this article can be effectively used to detect and identify dynamic NFTs, helping NFT holders avoid risks. Keting Yin, Xiaoxue Ren |
SMC | 3 |
| 2024 | $\mathbf{A^{3}}$A3-CodGen: A Repository-Level Code Generation Framework for Code Reuse With Local-Aware, Global-Aware, and Third-Party-Library-AwareabstractLLM-based code generation tools are essential to help developers in the software development process. Existing tools often disconnect with the working context, i.e., the code repository, causing the generated code to be not similar to human developers. In this paper, we propose a novel code generation framework, dubbed$A^{3}$-CodGen, to harness information within the code repository to generate code with fewer potential logical errors, code redundancy, and library-induced compatibility issues. We identify three types of representative information for the code repository: local-aware information from the current code file, global-aware information from other code files, and third-party-library information. Results demonstrate that by adopting the$A^{3}$-CodGen framework, we successfully extract, fuse, and feed code repository information into the LLM, generating more accurate, efficient, and highly reusable code. The effectiveness of our framework is further underscored by generating code with a higher reuse rate, compared to human developers. This research contributes significantly to the field of code generation, providing developers with a more powerful tool to address the evolving demands in software development in practice. Dianshu Liao, Shidong Pan, Xiaoyu Sun 0002, Xiaoxue Ren, Zhenchang Xing, Huan Jin, Qinying Li |
IEEE Trans. Software Eng. | 4 |
| 2023 | ConvMHSA-SCVD: Enhancing Smart Contract Vulnerability Detection through a Knowledge-Driven and Data-Driven FrameworkabstractSmart contracts are essential for executing computing logic on blockchain networks. However, they are also susceptible to various vulnerabilities. In recent years, the detection of smart contract vulnerabilities has become a significant concern due to the substantial losses caused by hacker attacks. Traditional vulnerability detection approaches rely on expert rules, which often suffer from limitations in accuracy and completeness. Deep learning-based methods offer better coverage of vulnerabilities but may overlook certain vulnerability characteristics and suffer from overfitting during training. In this paper, we propose a novel approach called ConvMHSA-SCVD, which combines knowledge-driven and data-driven algorithms together to detect smart contract vulnerabilities. By incorporating feature selection, data balancing, and a combination of multi-channel convolution and multi-head self-attention neural networks, our ConvMHSA-SCVD achieves effective vulnerability detection in smart contracts. Extensive experiments demonstrate that our approach outperforms the state-of-the-art method in accuracy and F1 score, with improvements ranging from 0.4% to 3.84% and 1.28% to 1.90%, respectively. Mengliang Li, Xiaoxue Ren, Zhuo Li 0014, Jianling Sun |
ISSRE | 2 |
| 2023 | From Misuse to Mastery: Enhancing Code Generation with Knowledge-Driven AI ChainingabstractLarge Language Models (LLMs) have shown promising results in automatic code generation by improving coding efficiency to a certain extent. However, generating high-quality and reliable code remains a formidable task because of LLMs' lack of good programming practice, especially in exception handling. In this paper, we first conduct an empirical study and summarize three crucial challenges of LLMs in exception handling, i.e., incomplete exception handling, incorrect exception handling and abuse of try-catch. We then try prompts with different granularities to address such challenges, finding fine-grained knowledge-driven prompts works best. Based on our empirical study, we propose a novel Knowledge-driven Prompt Chaining-based code generation approach, name KPC, which decomposes code generation into an AI chain with iterative check-rewrite steps and chains fine-grained knowledge-driven prompts to assist LLMs in considering exception-handling specifications. We evaluate our KPC-based approach with 3,079 code generation tasks extracted from the Java official API documentation. Extensive experimental results demonstrate that the KPC-based approach has considerable potential to ameliorate the quality of code generated by LLMs. It achieves this through proficiently managing exceptions and obtaining remarkable enhancements of 109.86% and 578.57% with static evaluation methods, as well as a reduction of 18 runtime bugs in the sampled dataset with dynamic validation. Xiaoxue Ren, Xinyuan Ye, Dehai Zhao, Zhenchang Xing, Xiaohu Yang 0001 |
ASE | 1 |
| 2023 | API-Knowledge Aware Search-Based Software Testing: Where, What, and HowabstractSearch-based software testing (SBST) has proved its effectiveness in generating test cases to achieve its defined test goals, such as branch and data-dependency coverage. However, to detect more program faults in an effective way, pre-defined goals can hardly be adaptive in diversified projects. In this work, we propose KAT, a novel knowledge-aware SBST approach to generate on-demand assertions in the program under test (PUT) based on its used APIs. KAT constructs an API knowledge graph from the API documentation to derive the constraints that the client codes need to satisfy. Each constraint is instrumented into the PUT as a program branch, serving as a test goal to guide SBST to detect faults. We evaluate KAT with two baselines (i.e., EvoSuite and Catcher) with a close-world and an open-world experiment to detect API bugs. The close-world experiment shows that KAT outperforms the baselines in the F1-score (0.55 vs. 0.24 and 0.30) to detect API-related bugs. The open-world experiment shows that KAT can detect 59.64% and 9.05% more bugs than the baselines in practice. Xiaoxue Ren, Xinyuan Ye, Yun Lin 0001, Zhenchang Xing, Shuqing Li 0001, Michael R. Lyu |
ESEC/SIGSOFT FSE | 1 |
| 2023 | The time-dependent electric vehicle routing problem with drone and synchronized mobile battery swapping
Xiaoxue Ren, Houming Fan, Ming-Xin Bao |
Adv. Eng. Informatics | 1 |
| 2023 | API Usage Recommendation Via Multi-View Heterogeneous Graph Representation LearningabstractDevelopers often need to decide which APIs to use for the functions being implemented. With the ever-growing number of APIs and libraries, it becomes increasingly difficult for developers to find appropriate APIs, indicating the necessity of automatic API usage recommendation. Previous studies adopt statistical models or collaborative filtering methods to mine the implicit API usage patterns for recommendation. However, they rely on the occurrence frequencies of APIs for mining usage patterns, thus prone to fail for the low-frequency APIs. Besides, prior studies generally regard the API call interaction graph as homogeneous graph, ignoring the rich information (e.g., edge types) in the structure graph. In this work, we propose a novel method namedMEGAfor improving the recommendation accuracy especially for the low-frequency APIs. Specifically, besidescall interaction graph, MEGA considers another two new heterogeneous graphs:global API co-occurrence graphenriched with the API frequency information andhierarchical structure graphenriched with the project component information. With the three multi-view heterogeneous graphs, MEGA can capture the API usage patterns more accurately. Experiments on three Java benchmark datasets demonstrate that MEGA significantly outperforms the baseline models by at least 19% with respect to the Success Rate@1 metric. Especially, for the low-frequency APIs, MEGA also increases the baselines by at least 55% regarding the Success Rate@1 score. Yujia Chen 0004, Cuiyun Gao 0001, Xiaoxue Ren, Yun Peng 0003, Xin Xia 0001, Michael R. Lyu |
IEEE Trans. Software Eng. | 3 |
| 2022 | Characterizing and Mitigating Anti-patterns of Alerts in Industrial Cloud SystemsabstractAlerts are crucial for requesting prompt human intervention upon cloud anomalies. The quality of alerts significantly affects the cloud reliability and the cloud provider’s business revenue. In practice, we observe on-call engineers being hindered from quickly locating and fixing faulty cloud services because of the vast existence of misleading, non-informative, non-actionable alerts. We call the ineffectiveness of alerts "anti-patterns of alerts". To better understand the anti-patterns of alerts and provide actionable measures to mitigate anti-patterns, in this paper, we conduct the first empirical study on the practices of mitigating anti-patterns of alerts in an industrial cloud system. We study the alert strategies and the alert processing procedure at Huawei Cloud, a leading cloud provider. Our study combines the quantitative analysis of millions of alerts in two years and a survey with eighteen experienced engineers. As a result, we summarized four individual anti-patterns and two collective anti-patterns of alerts. We also summarize four current reactions to mitigate the anti-patterns of alerts, and the general preventative guidelines for the configuration of alert strategy. Lastly, we propose to explore the automatic evaluation of the Quality of Alerts (QoA), including the indicativeness, precision, and handleability of alerts, as a future research direction that assists in the automatic detection of alerts’ anti-patterns. The findings of our study are valuable for optimizing cloud monitoring systems and improving the reliability of cloud services. Jiacheng Shen, Yuxin Su 0001, Xiaoxue Ren, Yongqiang Yang, Michael R. Lyu |
DSN | 4 |
| 2021 | Automating Developer Chat MiningabstractOnline chatrooms are gaining popularity as a communication channel between widely distributed developers of Open Source Software (OSS) projects. Most discussion threads in chatrooms follow a Q&A format, with some developers (askers) raising an initial question and others (respondents) joining in to provide answers. These discussion threads are embedded with rich information that can satisfy the diverse needs of various OSS stakeholders. However, retrieving information from threads is challenging as it requires a thread-level analysis to understand the context. Moreover, the chat data is transient and unstructured, consisting of entangled informal conversations. In this paper, we address this challenge by identifying the information types available in developer chats and further introducing an automated mining technique. Through manual examination of chat data from three chatrooms on Gitter, using card sorting, we build a thread-level taxonomy with nine information categories and create a labeled dataset with 2,959 threads. We propose a classification approach (named F2CHAT) to structure the vast amount of threads based on the information type automatically, helping stakeholders quickly acquire their desired information. F2CHAT effectively combines handcrafted non-textual features with deep textual features extracted by neural models. Specifically, it has two stages with the first one leveraging the siamese architecture to pretrain the textual feature encoder, and the second one facilitating an in-depth fusion of two types of features. Evaluation results suggest that our approach achieves an average F1-score of 0.628, which improves the baseline by 57%. Experiments also verify the effectiveness of our identified non-textual features under both intra-project and cross-project validations. Shengyi Pan, Lingfeng Bao, Xiaoxue Ren, Xin Xia 0001, David Lo 0001, Shanping Li |
ASE | 3 |
| 2021 | KGAMD: an API-misuse detector driven by fine-grained API-constraint knowledge graphabstractApplication Programming Interfaces (APIs) typically come with usage constraints. The violations of these constraints (i.e. API misuses) can cause significant problems in software development. Existing methods mine frequent API usage patterns from codebase to detect API misuses. They make a naive assumption that API usage that deviates from the most-frequent API usage is a misuse. However, there is a big knowledge gap between API usage patterns and API usage constraints in terms of comprehensiveness, explainability and best practices. Inspired by this, we propose a novel approach named KGAMD (API-Misuse Detector Driven by Fine-Grained API-Constraint Knowledge Graph) that detects API misuses directly against the API constraint knowledge, rather than API usage pat-terns. We first construct a novel API-constraint knowledge graph from API reference documentation with open information extraction methods. This knowledge graph explicitly models two types of API-constraint relations (call-order and condition-checking) and enriches return and throw relations with return conditions and exception triggers. Then, we develop the KGAMD tool that utilizes the knowledge graph to detect API misuses. There are three types of frequent API misuses we can detect - missing calls, missing condition checking and missing exception handling, while existing detectors mostly focus on only missing calls. Our quantitative evaluation and user study demonstrate that our KGAMD is promising in helping developers avoid and debug API misuses Xiaoxue Ren, Xinyuan Ye, Zhenchang Xing, Xin Xia 0001, Xiwei Xu 0001, Liming Zhu 0001, Jianling Sun |
ESEC/SIGSOFT FSE | 1 |
| 2020 | Demystify official API usage directives with crowdsourced API misuse scenarios, erroneous code examples and patchesabstractAPI usage directives in official API documentation describe the contracts, constraints and guidelines for using APIs in natural language. Through the investigation of API misuse scenarios on Stack Overflow, we identify three barriers that hinder the understanding of the API usage directives, i.e., lack of specific usage context, indirect relationships to cooperative APIs, and confusing APIs with subtle differences. To overcome these barriers, we develop a text mining approach to discover the crowdsourced API misuse scenarios on Stack Overflow and extract from these scenarios erroneous code examples and patches, as well as related API and confusing APIs to construct demystification reports to help developers understand the official API usage directives described in natural language. We apply our approach to API usage directives in official Android API documentation and android-tagged discussion threads on Stack Overflow. We extract 159,116 API misuse scenarios for 23,969 API usage directives of 3138 classes and 7471 methods, from which we generate the demystification reports. Our manual examination confirms that the extracted information in the generated demystification reports are of high accuracy. By a user study of 14 developers on 8 API-misuse related error scenarios, we show that our demystification reports help developer understand and debug API-misuse related program errors faster and more accurately, compared with reading only plain API usage-directive sentences. Xiaoxue Ren, Jiamou Sun, Zhenchang Xing, Xin Xia 0001, Jianling Sun |
ICSE | 1 |
| 2020 | API-Misuse Detection Driven by Fine-Grained API-Constraint Knowledge GraphabstractAPI misuses cause significant problem in software development. Existing methods detect API misuses against frequent API usage patterns mined from codebase. They make a naive assumption that API usage that deviates from the most-frequent API usage is a misuse. However, there is a big knowledge gap between API usage patterns and API usage caveats in terms of comprehensiveness, explainability and best practices. In this work, we propose a novel approach that detects API misuses directly against the API caveat knowledge, rather than API usage patterns. We develop open information extraction methods to construct a novel API-constraint knowledge graph from API reference documentation. This knowledge graph explicitly models two types of API-constraint relations (call-order and condition-checking) and enriches return and throw relations with return conditions and exception triggers. It empowers the detection of three types of frequent API misuses - missing calls, missing condition checking and missing exception handling, while existing detectors mostly focus on only missing calls. As a proof-of-concept, we apply our approach to Java SDK API Specification. Our evaluation confirms the high accuracy of the extracted API-constraint relations. Our knowledge-driven API misuse detector achieves 0.60 (68/113) precision and 0.28 (68/239) recall for detecting Java API misuses in the API misuse benchmark MuBench. This performance is significantly higher than that of existing pattern-based API misused detectors. A pilot user study with 12 developers shows that our knowledge-driven API misuse detection is very promising in helping developers avoid API misuses and debug the bugs caused by API misuses. Xiaoxue Ren, Xinyuan Ye, Zhenchang Xing, Xin Xia 0001, Xiwei Xu 0001, Liming Zhu 0001, Jianling Sun |
ASE | 1 |
| 2019 | Discovering, Explaining and Summarizing Controversial Discussions in Community Q&A SitesabstractDevelopers often look for solutions to programming problems in community Q&A sites like Stack Overflow. Due to the crowdsourcing nature of these Q&A sites, many user-provided answers are wrong, less optimal or out-of-date. Relying on community-curated quality indicators (e.g., accepted answer, answer vote) cannot reliably identify these answer problems. Such problematic answers are often criticized by other users. However, these critiques are not readily discoverable when reading the posts. In this paper, we consider the answers being criticized and their critique posts as controversial discussions in community Q&A sites. To help developers notice such controversial discussions and make more informed choices of appropriate solutions, we design an automatic open information extraction approach for systematically discovering and summarizing the controversies in Stack Overflow and exploiting official API documentation to assist the understanding of the discovered controversies. We apply our approach to millions of java/android-tagged Stack overflow questions and answers and discover a large scale of controversial discussions in Stack Overflow. Our manual evaluation confirms that the extracted controversy information is of high accuracy. A user study with 18 developers demonstrates the usefulness of our generated controversy summaries in helping developers avoid the controversial answers and choose more appropriate solutions to programming questions. Xiaoxue Ren, Zhenchang Xing, Xin Xia 0001, Guoqiang Li 0001, Jianling Sun |
ASE | 1 |
| 2019 | Neural Network-based Detection of Self-Admitted Technical Debt: From Performance to ExplainabilityabstractTechnical debt is a metaphor to reflect the tradeoff software engineers make between short-term benefits and long-term stability. Self-admitted technical debt (SATD), a variant of technical debt, has been proposed to identify debt that is intentionally introduced during software development, e.g., temporary fixes and workarounds. Previous studies have leveraged human-summarized patterns (which represent n-gram phrases that can be used to identify SATD) or text-mining techniques to detect SATD in source code comments. However, several characteristics of SATD features in code comments, such as vocabulary diversity, project uniqueness, length, and semantic variations, pose a big challenge to the accuracy of pattern or traditional text-mining-based SATD detection, especially for cross-project deployment. Furthermore, although traditional text-mining-based method outperforms pattern-based method in prediction accuracy, the text features it uses are less intuitive than human-summarized patterns, which makes the prediction results hard to explain. To improve the accuracy of SATD prediction, especially for cross-project prediction, we propose a Convolutional Neural Network-- (CNN) based approach for classifying code comments as SATD or non-SATD. To improve the explainability of our model’s prediction results, we exploit the computational structure of CNNs to identify key phrases and patterns in code comments that are most relevant to SATD. We have conducted an extensive set of experiments with 62,566 code comments from 10 open-source projects and a user study with 150 comments of another three projects. Our evaluation confirms the effectiveness of different aspects of our approach and its superior performance, generalizability, adaptability, and explainability over current state-of-the-art traditional text-mining-based methods for SATD classification. Xiaoxue Ren, Zhenchang Xing, Xin Xia 0001, David Lo 0001, Xinyu Wang 0001, John C. Grundy |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2018 | Characterizing Common and Domain-Specific Package Bugs: A Case Study on UbuntuabstractUbuntu is an open source software platform that runs everywhere from the smartphone, the tablet and the PC to the server and the cloud. In Ubuntu, there are many self-contained or third-party software packages for different use, and a bug report in Ubuntu could affect one or more packages simultaneously. Identifying the common package bugs in Ubuntu can help both developers and users better understand the packages they are developing or using, and also provide further guidelines to developers of similar packages in the future. In this paper, we perform a large-scale empirical study of common package bugs on Ubuntu by leveraging topic modeling. By analyzing a total of 240,097 bug reports, we identify 3 general bugs that are common to all Ubuntu packages, i.e., Graphical User Interface (GUI), Maintenance, and Runtime bugs. Moreover, we categorize top-100 packages with most number of bug reports into 6 categories (i.e., graphics, internet, office, sound and video, system management, and kernel), and identify domain-specific bugs for each category. Xiaoxue Ren, Xin Xia 0001, Zhenchang Xing, Lingfeng Bao, David Lo 0001 |
COMPSAC (1) | 1 |