Zhicao Tang

dblp:371/9430 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0003-0349-0850ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Commit Messages Generation Based on Core Changes
abstract
Commits messages play a crucial role in helping developers efficiently comprehend code modifications. Due to the time pressure of project iteration or poor message-writing practices, many commits suffer from missing messages. To address this issue, researchers have explored the automated generation of commit messages. Because of the truncation mechanism of the learning-based model, most of the current studies focus on code changes appearing at the beginning of a commit into the model for commit message generation. This may not be the best strategy for commit message generation because each code change in a commit contributes unequally to its overall purpose. To better generate commit messages, we propose a novel method that identifies the core code change in a commit for commit message generation. Specifically, we employ a method to predict the relative importance of the classes contained in a commit, and the code change of the class with the highest importance score (i.e., core change) is used to generate the commit message. Incorporating core change information can boost the performance of other existing methods (such as NMT, NNGen, and CoreGen). Building on this insight, we develop CCGen—a Core Change-Based Generation model that integrates a Transformer architecture with CodeBERT-enhanced encoding to leverage code semantics. The experiment demonstrates that the proposed method for commit message generation outperforms the state-of-the-art by 18.47% on average across seven metrics including 19.97 on ROUGE-L.
Yuan Huang 0002, Zhicao Tang, Xiangping Chen, Changlin Yang, Zibin Zheng, Xiaocong Zhou
ACM Trans. Softw. Eng. Methodol.2
2024 ESGen: Commit Message Generation Based on Edit Sequence of Code Change
abstract
Commit messages provide important information for comprehending the code changes, and a number of researchers try to generate commit messages by using an automatic way. These research on commit message generation has profited from the code tokens or code structures such as AST. Since the edit sequence of code change is also important for capturing the code change intent, we propose a new commit message generation method called ESGen, which extracts AST edit sequences of code changes as model input. Specifically, we employ an O(ND) difference algorithm to extract the edit sequence from AST by comparing the ASTs before and after applying the code changes. Then, we construct a Bi-Encoder, which encodes the textual information and the AST edit sequence information of code change. The experimental results show that ESGen outperforms other baseline models, improving the BLEU-4 to 15.14. Also, when applying the edit sequence to 7 baseline models, they improve the BLEU-4 scores of these models by an average of 8.5%. Additionally, a human evaluation confirmed the effectiveness of ESGen in generating commit messages.
Xiangping Chen, Yangzi Li, Zhicao Tang, Yuan Huang 0002, Haojie Zhou, Mingdong Tang, Zibin Zheng
ICPC3
2024 Towards automatically identifying the co-change of production and test code
abstract
Abstract In software evolution, keeping the test code co‐change with the production code is important, because the outdated test code may not work and is ineffective in revealing faults in the production code. However, due to the tight development time, the production and test code may not be co‐changed immediately by developers. For example, we analysed the top 1003 popular Java projects on GitHub and found that nearly 9.3% of cases (i.e., 464,417) did not update their production and test code at the same time, that is, the production code is updated first, and then the test code is updated at intervals. The result indicates that much test code will not be updated in time. In this paper, we propose a novel approach, Jtup, to remind developers to co‐change the production code and test code in time. Specifically, we first define the co‐changed production and test code as a positive instance, while unchanged test code (i.e., production code changed and test code unchanged) as a negative instance. Then, we extract multidimensional features from the production code to characterize the possibility of their co‐change, including code change features, code complexity features, and code semantic features. Finally, several machine learning‐based methods are employed to identify the co‐changed production and test code. We conduct comprehensive experiments on 20 datasets, and the results show that the Accuracy, Precision, and Recall achieved by Jtup are 76.7%, 78.1%, and 77.4%, which outperforms the state‐of‐the‐art method.
Yuan Huang 0002, Zhicao Tang, Xiangping Chen, Xiaocong Zhou
Softw. Test. Verification Reliab.2