VLDB 2026 Research / reviewers in the wild / expert
Chao Liu 0014
dblp:15/5923-14
· DBLP profile ↗
26ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-8283-9146ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 25 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UERR: A unified effective retrieval model for open-source repositories
Neng Zhang 0001, Jianga Shang, Haishen Lei, Chao Liu 0014, Yiwang Huang |
J. Syst. Softw. | 5 |
| 2026 | AdaCoder: An Adaptive Planning and Multi-Agent Framework for Function-Level Code GenerationabstractRecently, researchers have proposed many multi-agent frameworks for function-level code generation, which aim to improve software development productivity by automatically generating function-level source code based on task descriptions. A typical multi-agent framework consists of Large Language Model (LLM)-based agents that are responsible for task planning, code generation, testing, debugging, etc. Studies have shown that existing multi-agent code generation frameworks perform well on ChatGPT. However, their generalizability across other foundation LLMs remains unexplored systematically. In this paper, we report an empirical study on the generalizability of four state-of-the-art multi-agent code generation frameworks across 12 open-source LLMs with varying code generation and instruction-following capabilities. Our study reveals the unstable generalizability of existing frameworks on diverse foundation LLMs. Based on the findings obtained from the empirical study, we propose AdaCoder, a novel adaptive planning, multi-agent framework for function-level code generation. AdaCoder has two phases. Phase-1 is an initial code generation step without planning, which uses an LLM-based coding agent and a script-based testing agent to unleash LLM’s native power, identify cases beyond LLM’s power, and determine the errors hindering execution. Phase-2 adds a rule-based debugging agent and an LLM-based planning agent for iterative code generation with planning. Our evaluation shows that AdaCoder achieves higher generalizability on diverse LLMs. Compared to the best baseline MapCoder, AdaCoder is on average 27.69% higher in Pass@1, 16 times faster in inference, and 12 times lower in token consumption. Yueheng Zhu, Chao Liu 0014, Xiaoxue Ren, Zhongxin Liu 0002, Ruwei Pan, Hongyu Zhang 0002 |
IEEE Trans. Software Eng. | 2 |
| 2025 | An empirical study of ChatGPT-related projects and their issues on GitHubabstractDue to its powerful capabilities in natural language understanding and content generation, ChatGPT has received widespread attention since its launch in 2022. An increasing number of ChatGPT-related projects (that enhance the capabilities of ChatGPT, develop applications by calling ChatGPT APIs, etc.) are being released on GitHub and have sparked widespread discussions. However, GitHub does not provide a detailed classification of these projects to help users effectively explore interested projects. Additionally, the issues raised by users for these projects cover various aspects, e.g., installation, usage, and updates. It would be valuable to help developers prioritize more urgent issues and improve development efficiency. Unfortunately, there is currently no research focused on understanding the categories and issues of ChatGPT-related projects. To fill this gap, we retrieved 71,244 projects from GitHub using the keyword ‘ChatGPT’ and selected the top 200 representative projects with the highest numbers of stars as our dataset. By analyzing the project descriptions , we identified three primary categories of ChatGPT-related projects, namely ChatGPT Implementation & Training , ChatGPT Application , ChatGPT Improvement & Extension . We further built a classifier for automatically categorizing projects based on the 200 manually annotated projects. Next, we applied a topic modeling technique to 23,609 issues of those projects and identified ten issue topics, e.g., model reply and interaction interface . We analyzed the popularity, difficulty, and evolution of each issue topic within the three project categories and further proposed a method for recommending solutions for open issues by summarizing the pull requests associated with closed issues. Our main findings are: (1) The increase in the number of projects within the three categories is closely related to the development of ChatGPT; and (2) There are significant differences in the popularity, difficulty, and evolutionary trends of the issue topics across the three project categories. Based on these findings, we finally provided implications for project developers and platform managers on how to better develop and manage ChatGPT-related projects, such as offering more fine-grained tags to categorize projects to facilitate their exploration. Neng Zhang 0001, Chao Liu 0014, Zibin Zheng |
Expert Syst. Appl. | 3 |
| 2025 | GNPSum: A code summarization enhancement framework based on Graph Node Position
Haogang Cheng, Luwen Huangfu, Chao Liu 0014, Meng Yan 0001, Yan Lei 0005 |
Inf. Softw. Technol. | 4 |
| 2025 | DeepVec: State-Vector Aware Test Case Selection for Enhancing Recurrent Neural NetworkabstractDeep Neural Networks (DNN) have realized significant achievements across various application domains. There is no doubt that testing and enhancing a pre-trained DNN that has been deployed in an application scenario is crucial, because it can reduce the failures of the DNN. DNN-driven software testing and enhancement require large amounts of labeled data. The high cost and inefficiency caused by the large volume of data of manual labeling, and the time consumption of testing all cases in real scenarios are unacceptable. Therefore, test case selection technologies are proposed to reduce the time cost by selecting and only labeling representative test cases without compromising testing performance. Test case selection based on neuron coverage (NC) or uncertainty metrics has achieved significant success in Convolutional Neural Networks (CNN) testing. However, it is challenging to transfer these methods to Recurrent Neural Networks (RNN), which excel at text tasks, due to the mismatch in model output formats and the reliance on image-specific characteristics. What’s more, balancing the execution cost and performance of the algorithm is also indispensable.In this paper, we propose a state-vector aware test case selection method for RNN models, namely DeepVec, which reduces the cost of data labeling and saves computing resources and balances the execution cost and performance. DeepVec selects data using uncertainty metric based on the norm of the output vector at each time step (i.e., state-vector), and similarity metric based on the direction angle of the state-vector. Because test cases with smaller state-vector norms often possess greater information entropy and similar changes of state-vector direction angle indicate similar RNN internal states. These metrics can be calculated with just a single inference, which gives it strong bug detection and model improvement capabilities. We evaluate DeepVec on five popular datasets, containing images and texts as well as commonly used 3 RNN classification models, and compare it with NC-based, uncertainty-based, and other black-box methods. Experimental results demonstrate that DeepVec achieves an average relative improvement of 12.5%-118.22% over baseline methods in selecting fault-revealing test cases with time costs reduced to only 1% to 1‱. At the same time, we find that the absolute accuracy improvement after retraining outperforms baseline methods by 0.29%-24.01% when selecting 15% data to retrain. Zhonghao Jiang, Meng Yan 0001, Li Huang 0006, Weifeng Sun 0004, Chao Liu 0014, David Lo 0001 |
IEEE Trans. Software Eng. | 5 |
| 2025 | Improving Co-Decoding Based Security Hardening of Code LLMs Leveraging Knowledge DistillationabstractLarge Language Models (LLMs) have been widely adopted by developers in software development. However, the massive pretraining code data is not rigorously filtered, allowing LLMs to learn unsafe coding patterns. Several prior studies have demonstrated that code LLMs tend to generate code with potential vulnerabilities. The widespread adoption of intelligent programming assistants poses a significant threat to the software development process. Existing approaches to mitigating this risk primarily involve constructing secure data that are free of vulnerabilities and then retraining or fine-tuning the models. However, such an effort is resource intensive and requires significant manual supervision. When the model parameters are too large (e.g., more than 1 billion) or multiple models with the same parameter scale have the same optimization needs (e.g., to avoid outputting vulnerable code), the above work will become unaffordable. To address this challenge, in previous work, we proposed CoSec, an approach to improve the security of code LLMs with different parameters by utilizing an independent and very small parametric security model as a decoding navigator.Despite CoSec’s excellent performance, we found that there is still room for improving: 1) its ability to maintain the functional correctness of hardened targets, and 2) the security of the generated code. To address the above issues, we propose CoSec+, a hardening framework consisting of three phases: 1) Functional Correctness Alignment, which improves the functional correctness of the security base with knowledge disstillation; 2) Security Training, which yields an independent, but much smaller security model; and 3) Co-decoding, where the security model iteratively reasons about the next token along with the target model. Due to the higher confidence that a well-trained security model places in secure and correct tokens, it guides the target base model to generate more secure code, even as it improves the functional correctness of the target base model. We have conducted extensive experiments in several code LLMs (i.e., CodeGen, StarCoderBase, DeepSeekCoder and Qwen2.5-Coder), and the results show that our approach is effective in improving the functional correctness and security of the models. The evaluation results show that CoSec+ can deliver a 0.8% to 37.7% improvement in security across models of various parameter sizes and families; moreover, it preserves the functional correctness of the target base models—achieving functional-correctness gains of 0.7% to 51.1% for most of those models. Dong Li 0009, Shanfu Shu, Meng Yan 0001, Zhongxin Liu 0002, Chao Liu 0014, Xiaohong Zhang 0002, David Lo 0001 |
IEEE Trans. Software Eng. | 5 |
| 2025 | Hydra-Reviewer: A Holistic Multi-Agent System for Automatic Code Review Comment GenerationabstractReview comment generation is a crucial task in code review, and significant progress has been made in automating. Previous research has generated review comments by fine-tuning pre-trained models or Large Language Models (LLMs). However, these studies have overlooked the necessity of conducting code reviews from multiple perspectives, resulting in the omission of potential issues in code changes. Additionally, the complexity of review comments often hinders the accurate quantitative evaluation of automated tools’ effectiveness.In this paper, we first conduct an empirical study to propose a comprehensive taxonomy of code review dimensions. We also identify three major limitations of existing automated code review (ACR) methods: lack of comprehensiveness, incorrectness, and vagueness. Building on the insights from our empirical study, we introduce Hydra-Reviewer, a collaborative multi-agent framework powered by large language models, designed to automatically generate high-quality code reviews. We utilize the CodeReview and CodeReviewNew benchmark datasets, along with a newly constructed review comment generation dataset. We compare Hydra-Reviewerwith several baselines, including CodeReviewer, LLaMA-Reviewer, ChatGPT, Comprehensive-ChatGPT, and DeepSeek-V3.The experimental results show that Hydra-Reviewerachieves a BLEU score of 8.20, outperforming the state-of-the-art baseline, DeepSeek-V3, which scores 7.85. In qualitative evaluation, Hydra-Reviewer’s generated comments span an average of 7.8 review dimensions, addressing the limitations of existing ACR methods effectively. Additionally, Hydra-Reviewerdemonstrates strong generalization capabilities on unseen dataset. We further validate the contributions of each component of Hydra-Reviewerthrough an ablation study and confirm the helpfulness and readability of the generated comments via a User Study. Finally, a cost analysis reveals that Hydra-Reviewergenerates review comments at an average cost of 0.018 dollars and 62.63 seconds per code change. Xiaoxue Ren, Chaoqun Dai, Ye Wang 0012, Chao Liu 0014, Bo Jiang 0009 |
IEEE Trans. Software Eng. | 5 |
| 2024 | VisRepo: A Visual Retrieval Tool for Large-Scale Open-Source ProjectsabstractTo improve software development productivity, developers frequently search for projects on open-source communities such as GitHub. However, it is challenging for users to quickly find suitable projects from numerous results due to the overload of project information. Although many tools have been proposed to rank the relevancy of searched results, manually inspecting them one by one is irreplaceable and time-consuming. To fill this gap, we propose a visual retrieval tool named VisRepo for open-source software projects. Firstly, it mines software project data from four perspectives including topic, technology, usability, and comprehensibility, and connects projects based on the same owners/contributors and similar topics. Then, visualization technique is employed to present complex software data intuitively. VisRepo provides users an interactive retrieval paradigm of Search-Explore-Check-Recommend with in-depth insights and better exploration experience. We evaluate VisRepo on 7w+ open-source JavaScript projects. Experimental results showed that VisRepo outperforms GitHub search engine in terms of time consumption and accuracy, meanwhile enabling a more interactive and useful user experience. Xiaoqi Yue, Chao Liu 0014, Neng Zhang 0001, Haibo Hu 0002, Xiaohong Zhang 0002 |
Internetware | 2 |
| 2024 | CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingabstractLarge Language Models (LLMs) specialized in code have shown exceptional proficiency across various programming-related tasks, particularly code generation. Nonetheless, due to its nature of pretraining on massive uncritically filtered data, prior studies have shown that code LLMs are prone to generate code with potential vulnerabilities. Existing approaches to mitigate this risk involve crafting data without vulnerability and subsequently retraining or fine-tuning the model. As the number of parameters exceeds a billion, the computation and data demands of the above approaches will be enormous. Moreover, an increasing number of code LLMs tend to be distributed as services, where the internal representation is not accessible, and the API is the only way to reach the LLM, making the prior mitigation strategies non-applicable. To cope with this, we propose CoSec, an on-the-fly Security hardening method of code LLMs based on security model-guided Co-decoding, to reduce the likelihood of code LLMs to generate code containing vulnerabilities. Our key idea is to train a separate but much smaller security model to co-decode with a target code LLM. Since the trained secure model has higher confidence for secure tokens, it guides the generation of the target base model towards more secure code generation. By adjusting the probability distributions of tokens during each step of the decoding process, our approach effectively influences the tendencies of generation without accessing the internal parameters of the target code LLM. We have conducted extensive experiments across various parameters in multiple code LLMs (i.e., CodeGen, StarCoder, and DeepSeek-Coder), and the results show that our approach is effective in security hardening. Specifically, our approach improves the average security ratio of six base models by 5.02%-37.14%, while maintaining the functional correctness of the target model. Dong Li 0009, Meng Yan 0001, Yaosheng Zhang, Zhongxin Liu 0002, Chao Liu 0014, Xiaohong Zhang 0002, Ting Chen 0002, David Lo 0001 |
ISSTA | 5 |
| 2024 | Guiding ChatGPT for Better Code Generation: An Empirical StudyabstractAutomated code generation is a powerful technique for software development, which can significantly reduce developers' effort and time for writing code. Recently, OpenAI's large language model ChatGPT has emerged as a powerful tool for generating human-like responses to a wide range of textual inputs (i.e., prompts), including those related to code generation. However, the effectiveness of ChatGPT in code generation is still not well understood. The code generation performance could also be heavily influenced by the choice of prompts, which should be further explored. In this paper, we report an empirical study on ChatGPT's capabilities for two types of code generation tasks, namely text-to-code and code-to-code generation. We investigate different types of prompts by leveraging the chain-of-thought strategy with multi-step optimizations. Our empirical results show that by carefully designing prompts to guide ChatGPT, the code generation performance can be improved substantially. We also analyze the factors that influence the prompt design and provide insights that could guide future research. Chao Liu 0014, Xuanlin Bao, Hongyu Zhang 0002, Neng Zhang 0001, Haibo Hu 0002, Xiaohong Zhang 0002, Meng Yan 0001 |
SANER | 1 |
| 2024 | Understanding the implementation issues when using deep learning frameworks
Chao Liu 0014, Runfeng Cai, Yiqun Zhou, Haibo Hu 0002, Meng Yan 0001 |
Inf. Softw. Technol. | 1 |
| 2024 | Code semantic enrichment for deep code search
Zhongyang Deng, Chao Liu 0014, Luwen Huangfu, Meng Yan 0001 |
J. Syst. Softw. | 3 |
| 2024 | End-to-end log statement generation at block-level
Meng Yan 0001, Pinjia He, Chao Liu 0014, Xiaohong Zhang 0002, Dan Yang 0001 |
J. Syst. Softw. | 4 |
| 2024 | Query-oriented two-stage attention-based model for code search
Huanhuan Yang, Chao Liu 0014, Luwen Huangfu |
J. Syst. Softw. | 3 |
| 2022 | ShellFusion: Answer Generation for Shell Programming Tasks via Knowledge FusionabstractShell commands are widely used for accomplishing tasks, such as network management and file manipulation, in Unix and Linux platforms. There are a large number of shell commands available. For example, 50,000+ commands are documented in the Ubuntu Manual Pages (MPs). Quite often, programmers feel frustrated when searching and orchestrating appropriate shell commands to accomplish specific tasks. To address the challenge, the shell programming community calls for easy-to-use tutorials for shell commands. However, existing tutorials (e.g., TLDR) only cover a limited number of frequently used commands for shell beginners and provide limited support for users to search for commands by a task. Neng Zhang 0001, Chao Liu 0014, Xin Xia 0001, Christoph Treude, Ying Zou 0001, David Lo 0001, Zibin Zheng |
ICSE | 2 |
| 2022 | CodeMatcher: a tool for large-scale code search based on query semantics matchingabstractDue to the emergence of large-scale codebases, such as GitHub and Gitee, searching and reusing existing code can help developers substantially improve software development productivity. Over the years, many code search tools have been developed. Early tools leveraged the information retrieval (IR) technique to perform an efficient code search for a frequently changed large-scale codebase. However, the search accuracy was low due to the semantic mismatch between query and code. In the recent years, many tools leveraged Deep Learning (DL) technique to address this issue. But the DL-based tools are slow and the search accuracy is unstable. Chao Liu 0014, Xuanlin Bao, Xin Xia 0001, Meng Yan 0001, David Lo 0001, Ting Zhang 0011 |
ESEC/SIGSOFT FSE | 1 |
| 2022 | Fine-grained Co-Attentive Representation Learning for Semantic Code SearchabstractCode search aims to find code snippets from large-scale code repositories based on the developer's query intent. A significant challenge for code search is the semantic gap between programming language and natural language. Recent works have indicated that deep learning (DL) techniques can perform well by automatically learning the relationships between query and code. Among these DL-based approaches, the state-of-the-art model is TabCS, a two-stage attention-based model for code search. However, TabCS still has two limitations: semantic loss and semantic confusion. TabCS breaks the structural information of code into token-level words of abstract syntax tree (AST), which loses the sequential semantics between words in programming statements, and it uses a co-attention mechanism to build the semantic correlation of code-query after fusing all features, which may confuse the correlations between individual code features and query. In this paper, we propose a code search model named FcarCS (Fine-grained Co-Attentive Representation Learning Model for Semantic Code Search). FcarCS extracts code textual features (i.e., method name, API sequence, and tokens) and structural features that introduce a statement-level code structure. Unlike TabCS, FcarCS splits AST into a series of subtrees corresponding to code statements and treats each subtree as a whole to preserve sequential semantics between words in code statements. FcarCS constructs a new fine-grained co-attention mechanism to learn interdependent representations for each code feature and query, respectively, instead of performing one co-attention process for the fused code features like TabCS. Generally, this mechanism leverages row/column-wise CNN to enable our model to focus on the strongly correlated local information between code feature and Query. We train and evaluate FcarCS on an open Java dataset with 475k and 10k code/query pairs, respectively. Experimental results show that FcarCS achieves an MRR of 0.613, outperforming three state-of-the-art models DeepCS, UNIF, and TabCS, by 117.38%, 16.76%, and 12.68%, respectively. We also performed a user study for each model with 50 real-world queries, and the results show that FcarCS returned code snippets that are more relevant than the baseline models. Zhongyang Deng, Chao Liu 0014, Meng Yan 0001, Zhou Xu 0003, Yan Lei 0005 |
SANER | 3 |
| 2022 | On the Reproducibility and Replicability of Deep Learning in Software EngineeringabstractContext:Deep learning (DL) techniques have gained significant popularity among software engineering (SE) researchers in recent years. This is because they can often solve many SE challenges without enormous manual feature engineering effort and complex domain knowledge. Objective:Although many DL studies have reported substantial advantages over other state-of-the-art models on effectiveness, they often ignore two factors:(1) reproducibility—whether the reported experimental results can be obtained by other researchers using authors’ artifacts (i.e., source code and datasets) with the same experimental setup; and(2) replicability—whether the reported experimental result can be obtained by other researchers using their re-implemented artifacts with a different experimental setup. We observed that DL studies commonly overlook these two factors and declare them as minor threats or leave them for future work. This is mainly due to high model complexity with many manually set parameters and the time-consuming optimization process, unlike classical supervised machine learning (ML) methods (e.g., random forest). This study aims to investigate the urgency and importance of reproducibility and replicability for DL studies on SE tasks. Method:In this study, we conducted a literature review on 147 DL studies recently published in 20 SE venues and 20 AI (Artificial Intelligence) venues to investigate these issues. We also re-ran four representative DL models in SE to investigate important factors that may strongly affect the reproducibility and replicability of a study. Results:Our statistics show the urgency of investigating these two factors in SE, where only 10.2% of the studies investigate any research question to show that their models can address at least one issue of replicability and/or reproducibility. More than 62.6% of the studies do not even share high-quality source code or complete data to support the reproducibility of their complex models. Meanwhile, our experimental results show the importance of reproducibility and replicability, where the reported performance of a DL model could not be reproduced for an unstable optimization process. Replicability could be substantially compromised if the model training is not convergent, or if performance is sensitive to the size of vocabulary and testing data. Conclusion:It is urgent for the SE community to provide a long-lasting link to a high-quality reproduction package, enhance DL-based solution stability and convergence, and avoid performance sensitivity on different sampled data. Chao Liu 0014, Cuiyun Gao 0001, Xin Xia 0001, David Lo 0001, John C. Grundy, Xiaohu Yang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2022 | CodeMatcher: Searching Code Based on Sequential Semantics of Important Query WordsabstractTo accelerate software development, developers frequently search and reuse existing code snippets from a large-scale codebase, e.g., GitHub. Over the years, researchers proposed many information retrieval (IR)-based models for code search, but they fail to connect the semantic gap between query and code. An early successful deep learning (DL)-based model DeepCS solved this issue by learning the relationship between pairs of code methods and corresponding natural language descriptions. Two major advantages of DeepCS are the capability of understanding irrelevant/noisy keywords and capturing sequential relationships between words in query and code. In this article, we proposed an IR-based model CodeMatcher that inherits the advantages of DeepCS (i.e., the capability of understanding the sequential semantics in important query words), while it can leverage the indexing technique in the IR-based model to accelerate the search response time substantially. CodeMatcher first collects metadata for query words to identify irrelevant/noisy ones, then iteratively performs fuzzy search with important query words on the codebase that is indexed by the Elasticsearch tool and finally reranks a set of returned candidate code according to how the tokens in the candidate code snippet sequentially matched the important words in a query. We verified its effectiveness on a large-scale codebase with ~41K repositories. Experimental results showed that CodeMatcher achieves an MRR (a widely used accuracy measure for code search) of 0.60, outperforming DeepCS, CodeHow, and UNIF by 82%, 62%, and 46%, respectively. Our proposed model is over 1.2K times faster than DeepCS. Moreover, CodeMatcher outperforms two existing online search engines (GitHub and Google search) by 46% and 33%, respectively, in terms of MRR. We also observed that: fusing the advantages of IR-based and DL-based models is promising; improving the quality of method naming helps code search, since method name plays an important role in connecting query and code. Chao Liu 0014, Xin Xia 0001, David Lo 0001, Ahmed E. Hassan, Shanping Li |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2022 | Multi-Dimension Convolutional Neural Network for Bug LocalizationabstractSoftware bugs remain frequent in the life cycle of software development and maintenance. Automatic localization of buggy source code files is critical for timely bug fixing and improving the efficiency of software quality assurance. Various bug localization techniques have been proposed using different dimensions of features. Recent studies have shown that different dimensions of features may play different roles in bug localization. Unfortunately, how to effectively merge these dimensions of features for improving bug localization has rarely been investigated. This article presents a Multi-Dimension Convolutional Neural Network (MD-CNN) model for bug localization automatically based on a bug report. Our approach has dual-novelty. First, we identify and extract five statistical dimensions of features. Second, we design a Convolutional Neural Network (CNN) model that takes our five statistical dimensions of features as the input and iteratively learns the complex and non-linear relationship between the features and the bug locations. The MD-CNN bug localization model is verified using six large-scale open source projects. The experimental results show that our MD-CNN outperforms the existing representative bug localization techniques in terms of the Mean Average Precision (MAP) and the number of bugs successfully localized in the top 1, 5, and 10 matched source code files. Bei Wang 0010, Meng Yan 0001, Chao Liu 0014, Ling Liu 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Two-Stage Attention-Based Model for Code Search with Textual and Structural FeaturesabstractSearching and reusing existing code from a large scale codebase can largely improve developers’ programming efficiency. To support code reuse, early code search models leverage information retrieval (IR) techniques to index a large-scale code corpus and return relevant code according to developers’ search query. However, IR-based models fail to capture the semantics in code and query. To tackle this issue, developers applied deep learning (DL) techniques to code search models. However, these models either are too complex to determine an effective method efficiently or learning for semantic correlation between code and query inadequately.To bridge the semantic gap between code and query effectively and efficiently, we propose a code search model TabCS (Two-stage Attention-Based model for Code Search) in this study. TabCS extracts code and query information from the code textual features (i.e., method name, API sequence, and tokens), the code structural feature (i.e., abstract syntax tree), and the query feature (i.e., tokens). TabCS performs a two-stage attention net-work structure. The first stage leverages attention mechanisms to extract semantics from code and query considering their semantic gap. The second stage leverages a co-attention mechanism to capture their semantic correlation and learn better code/query representation. We evaluate the performance of TabCS on two existing large-scale datasets with 485k and 542k code snippets, respectively. Experimental results show that TabCS achieves an MRR of 0.57 on Hu et al.’s dataset, outperforming three state-of-the-art models CARLCS-CNN, DeepCS, and UNIF by 18%, 70%, 12%, respectively. Meanwhile, TabCS gains an MRR of 0.54 on Husain et al.’s, outperforming CARLCS-CNN, DeepCS, and UNIF by 32%, 76%, 29%, respectively. Huanhuan Yang, Chao Liu 0014, Jianhang Shuai, Meng Yan 0001, Yan Lei 0005, Zhou Xu 0003 |
SANER | 3 |
| 2020 | Improving Code Search with Co-Attentive Representation LearningabstractSearching and reusing existing code from a large-scale codebase, e.g, GitHub, can help developers complete a programming task efficiently. Recently, Gu et al. proposed a deep learning-based model (i.e., DeepCS), which significantly outperformed prior models. The DeepCS embedded codebase and natural language queries into vectors by two LSTM (long and short-term memory) models separately, and returned developers the code with higher similarity to a code search query. However, such embedding method learned two isolated representations for code and query but ignored their internal semantic correlations. As a result, the learned isolated representations of code and query may limit the effectiveness of code search. Jianhang Shuai, Chao Liu 0014, Meng Yan 0001, Xin Xia 0001, Yan Lei 0005 |
ICPC | 3 |
| 2019 | A two-phase transfer learning model for cross-project defect predictionabstractContext: Previous studies have shown that a transfer learning model, TCA+ proposed by Nam et al., can significantly improve the performance of cross-project defect prediction (CPDP). TCA+ achieves the improvement by reducing data distribution difference between source (training data) and target (testing data) projects. However, TCA+ is unstable, i.e., its performance varies largely when using different source projects to build prediction models. In practice, it is hard to choose a suitable source project to build the prediction model. Objective: To address the limitation of TCA+, we propose a two-phase transfer learning model (TPTL) for CPDP. Method: In the first phase, we propose a source project estimator (SPE) to automatically choose two source projects with the highest distribution similarity to a target project from candidates. Next, two source projects that are estimated to achieve the highest values of F1-score and cost-effectiveness are selected. In the second phase, we leverage TCA+ to build two prediction models based on the two selected projects and combine their prediction results to further improve the prediction performance. Results: We evaluate TPTL on 42 defect datasets from PROMISE repository, and compare it with two versions of TCA+ (TCA+_Rnd, randomly selecting one source project; TCA+_All, using all alternative source projects), a related source project selection model TDS proposed by Herbold, a state-of-the-art CPDP model leveraging a log transformation (LT) method, and a transfer learning model Dycom with better form of TCA. Experiment results show that, on average across 42 datasets, TPTL respectively improves these baseline models by 19%, 5%, 36%, 27%, and 11% in terms of F1-score; by 64%, 92%, 71%, 11%, and 66% in terms of cost-effectiveness. Conclusion: The proposed TPTL model can solve the instability problem of TCA+, showing substantial improvements over the state-of-the-art and related CPDP models. Chao Liu 0014, Dan Yang 0001, Xin Xia 0001, Meng Yan 0001, Xiaohong Zhang 0002 |
Inf. Softw. Technol. | 1 |
| 2018 | Cross-Project Change-Proneness PredictionabstractSoftware change-proneness prediction (whether or not class files in a project will be changed in the next release) can help software developers to focus on preventive actions to reduce maintenance costs, and managers to allocate resources more effectively. Prior studies found that change-proneness prediction works well if there is sufficient amount of training data to build a model. However, it is not feasible for projects with limited historical data especially for new projects. To address this issue, cross-project change-proneness prediction, which builds a prediction model by using data in another project (i.e., source project), and predicts the change-proneness in a target project, is proposed. Considering there are a large number of source projects, one challenge for cross-project change-proneness prediction is that given a target project, how to automatically select a source project which could show good prediction accuracy on it. In this paper, we propose a selective cross-project (SCP) model for change-proneness prediction. SCP automatically finds the source project which has the similar data distribution with the target project by measuring distribution similarity between source and target projects. We evaluate SCP by conducting an empirical study on 14 open source projects. We compare it with 2 most related change-proneness models, including RCP (Random Cross-Project prediction) proposed by Malhotra and Bansal, and CLAMI+ developed by Yan et al. Experiment results show that SCP improves RCP and CLAMI+ by 25.34% and 4.30% in terms of AUC respectively; and by 171.42% and 172.31% in terms of cost-effectiveness, respectively. Chao Liu 0014, Dan Yang 0001, Xin Xia 0001, Meng Yan 0001, Xiaohong Zhang 0002 |
COMPSAC (1) | 1 |
| 2017 | Automated change-prone class prediction on unlabeled dataset using unsupervised method
Meng Yan 0001, Xiaohong Zhang 0002, Chao Liu 0014, Mengning Yang, Dan Yang 0001 |
Inf. Softw. Technol. | 3 |
| 2016 | Self-learning Change-prone Class PredictionabstractSoftware change-prone class prediction can enhance software decision making activities during software maintenance (e.g., resource allocating).Many change-prone class prediction approaches have been proposed and most are effective in interversion prediction within a project.These approaches usually build a supervised prediction model by learning from historical labeled dataset.However, a major challenge which remains is that this typical change-prone prediction setting cannot be used for new projects or projects with limited historical data.To address this challenge, we propose to tackle this task by adopting a novel prediction method which has not been used in changeprone prediction, namely self-learning method.The key idea of the self-learning method is to enable the change-prone prediction on new projects or projects with limited historical dataset by learning from itself.In this paper, we apply a state-of-art selflearning method, CLAMI, to change-prone prediction.In addition, we propose a novel self-learning approach CLAMI+ by extending CLAMI.The experiments among 14 open source projects show that the self-learning methods achieve comparable results to four typical inter-version baselines and the proposed CLAMI+ slightly improves the CLAMI method on average. Meng Yan 0001, Mengning Yang, Chao Liu 0014, Xiaohong Zhang 0002 |
SEKE | 3 |