Zhiheng Qu

dblp:325/0338 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Modeling Structural Features with Structure Position-Aware Attention for Code Summarization
abstract
The automatic generation of code comments is a crucial aspect in software engineering as it allows the description of source code functions in natural language. In recent years, researchers have utilized advanced machine translation models, particularly transformer-based models, to achieve remarkable results in the task of code summarization. Despite these advancements, current methods still face some challenges. First, existing models fail to incorporate the rich structural information of the code well, resulting in inadequate summary statements. Second, most generative models are too sensitive to the editing of code text, which results in insufficient generalization ability of the model. To address these issues, we proposed a new model named SPA-Trans(Structure Position-aware Attention Transformer-based model). SPA-Trans uses a distance matrix based on the relative distance of nodes in the AST (Abstract Syntax Tree) to represent the structural correlation between each token, enhancing the model’s ability to capture structural information of the source code. Additionally, in order to reduce the sensitivity of the model to code editing, we used the adversarial training method in the embedding layer to simulate code editing, which improves the generalization of the model. Our experiments on real-world datasets in Java and Python validate the effectiveness of our proposed method.
Zhiheng Qu, Bo Cai 0003
Int. J. Softw. Eng. Knowl. Eng.2
2022 Method Name Generation Based on Code Structure Guidance
abstract
The proper names of software engineering functions and methods can greatly assist developers in understanding and maintaining the code. Most researchers convert the method name generation task into the text summarization task. They take the token sequence and the abstract syntax tree (AST) of source code as input, and generate method names with a decoder. However, most proposed models learn semantic and structural features of the source code separately, resulting in poor performance in the method name generation task. Actually, each token in source code must have a corresponding node in its AST. Inspired by this observation, we propose SGMNG, a structure-guided method name generation model that learns the representation of two combined features. Additionally, we build a code graph called code relation graph (CRG) to describe the code structure clearly. CRG retains the structure of the AST of source code and contains data flows and control flows. SGMNG captures the semantic features of the code by encoding the token sequence and captures the structural features of the code by encoding the CRG. Then, SGMNG matches tokens in the sequence and nodes in the CRG to construct the combination of two features. We demonstrate the effectiveness of the proposed approach on the public dataset Java-Small with 700K samples, which indicates that our approach achieves significant improvement over the state-of-the-art baseline models in the ROUGE metric.
Zhiheng Qu, Jianhui Zeng, Bo Cai 0003
SANER1