Zijie Huang 0001

dblp:246/8147-1 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0002-8911-9889ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Usage patterns of software product metrics in assessing developers' output: A comprehensive study
Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Yuguo Liang
Inf. Softw. Technol.4
2026 Revisiting pre-trained models and feature fusion strategies for just-in-time defect prediction
Yuguo Liang, Guisheng Fan, Huiqun Yu, Chengcheng Wu, Zijie Huang 0001
Inf. Softw. Technol.6
2026 Automatic identification of extrinsic bug reports for just-in-time bug prediction
Guisheng Fan, Yuguo Liang, Longfei Zu, Huiqun Yu, Zijie Huang 0001
Sci. Comput. Program.5
2026 Software vulnerability detection via multimodal retrieval and hierarchical decision making
Zhoufu Liu, Zijie Huang 0001
Sci. Comput. Program.3
2025 Tool or Toy: Are SCA tools ready for challenging scenarios?
Congyan Shu, Guisheng Fan, Huiqun Yu, Zijie Huang 0001, Yuguo Liang
Comput. Secur.5
2025 Automatic Code Summarization Using Abbreviation Expansion and Subword Segmentation
abstract
ABSTRACT Automatic code summarization refers to generating concise natural language descriptions for code snippets. It is vital for improving the efficiency of program understanding among software developers and maintainers. Despite the impressive strides made by deep learning‐based methods, limitations still exist in their ability to understand and model semantic information due to the unique nature of programming languages. We propose two methods to boost code summarization models: context‐based abbreviation expansion and unigram language model‐based subword segmentation. We use heuristics to expand abbreviations within identifiers, reducing semantic ambiguity and improving the language alignment of code summarization models. Furthermore, we leverage subword segmentation to tokenize code into finer subword sequences, providing more semantic information during training and inference, thereby enhancing program understanding. These methods are model‐agnostic and can be readily integrated into existing automatic code summarization approaches. Experiments conducted on two widely used Java code summarization datasets demonstrated the effectiveness of our approach. Specifically, by fusing original and modified code representations into the Transformer model, our Semantic Enhanced Transformer for Code Summarizsation (SETCS) serves as a robust semantic‐level baseline. By simply modifying the datasets, our methods achieved performance improvements of up to 7.3%, 10.0%, 6.7%, and 3.2% for representative code summarization models in terms of BLEU‐4 , METEOR , ROUGE‐L and SIDE , respectively.
Yuguo Liang, Guisheng Fan, Huiqun Yu, Zijie Huang 0001
Expert Syst. J. Knowl. Eng.5
2025 RESEARCH NOTES: Multiclass Classification for Self-Admitted Technical Debt via Large Pre-Trained Language Model
abstract
Technical debt refers to suboptimal solutions adopted for short-term goals. Self-admitted technical debt (SATD) is the debt that is explicitly marked through comments or documentation, making it traceable. Multi-classification of SATD helps developers understand different debt types and improve efficiency. This paper proposes a SATD multi-classification method based on Fine-Tuning the GPT-3.5-turbo model for SATD prediction. This study uses a public dataset containing 10 projects with code comments. We classify design debt, requirement debt, and defect debt and evaluate our method’s performance. The experimental results show that compared to the best baseline model, our method achieves average improvements of 11.41%, 1.72% and 3.72% in MacroF, MacroP and MacroR metrics, respectively, in the MTO scenario. In the OTO scenario, improvements are 2.33%, 3.70% and 2.18%, respectively. These results indicate that our method has a strong generalization ability in SATD multi-classification and offers a new approach to managing technical debt.
Yiyang Du, Xingguang Yang, Zhenyu Shu, Zijie Huang 0001, Gang Wang 0023, Libo Xu
Int. J. Softw. Eng. Knowl. Eng.4
2024 A code change-oriented approach to just-in-time defect prediction with multiple input semantic fusion
abstract
Abstract Recent research found that fine‐tuning pre‐trained models is superior to training models from scratch in just‐in‐time (JIT) defect prediction. However, existing approaches using pre‐trained models have their limitations. First, the input length is constrained by the pre‐trained models.Secondly, the inputs are change‐agnostic.To address these limitations, we propose JIT‐Block, a JIT defect prediction method that combines multiple input semantics using changed block as the fundamental unit. We restructure the JIT‐Defects4J dataset used in previous research. We then conducted a comprehensive comparison using eleven performance metrics, including both effort‐aware and effort‐agnostic measures, against six state‐of‐the‐art baseline models. The results demonstrate that on the JIT defect prediction task, our approach outperforms the baseline models in all six metrics, showing improvements ranging from 1.5% to 800% in effort‐agnostic metrics and 0.3% to 57% in effort‐aware metrics. For the JIT defect code line localization task, our approach outperforms the baseline models in three out of five metrics, showing improvements of 11% to 140%.
Teng Huang 0004, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Chen-Yu Wu
Expert Syst. J. Knowl. Eng.4
2024 Aligning XAI explanations with software developers' expectations: A case study with code smell prioritization
Zijie Huang 0001, Huiqun Yu, Guisheng Fan, Zhiqing Shao, Yuguo Liang
Expert Syst. Appl.1
2024 Enhancing code summarization with action word prediction
Huiqun Yu, Guisheng Fan, Ziyi Zhou 0002, Zijie Huang 0001
Neurocomputing5
2024 On the effectiveness of developer features in code smell prioritization: A replication study
Zijie Huang 0001, Huiqun Yu, Guisheng Fan, Zhiqing Shao, Ziyi Zhou 0002
J. Syst. Softw.1
2024 Bug report priority prediction using social and technical features
abstract
Summary Software stakeholders report bugs in issue tracking system (ITS) with manually labeled priorities. However, the lack of knowledge and standard for prioritization may cause stakeholders to mislabel the priorities. In response, priority predictors are actively developed to support them. Prior studies trained machine learners based on textual similarity, categorical, and numeric technical features of bug reports. Most models were validated by time‐insensitive approaches, and they were producing suboptimal results for practical usage. While they ignored the social aspects of ITS, the technical aspects were also limited in surface features of bug reports. To better model the bug report, we extract their topic and most similar code structures. Since ITS bridges users and developers as the main contributors, we also integrate their experience, sentiment, and socio‐technical features to construct a new dataset. Then, we perform two‐classed and multiclassed bug priority prediction based on the dataset. We also introduce adversarial training using generated training data with random word swap and random word deletion. We validate our model in within‐project, cross‐project, and time‐wise scenarios, and it outperforms the two baselines by up to 15% in area under curve‐receiver operating characteristics (AUC‐ROC) and 19% in Matthews correlation coefficient (MCC). We reveal involving contributor (i.e., assignee and reporter) features such as sentiment that could boost prediction performance. Finally, we test statistically the mean and distribution of the features that reflect the differences in social and technical aspects (e.g., quality of communication and resource distribution) between high and low priority reports. In conclusion, we suggest that researchers should consider both social and technical aspects of ITS in bug report priority prediction and introduce adversarial training to boost model performance.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
J. Softw. Evol. Process.1
2024 Exploring better alternatives to size metrics for explainable software defect prediction
Chenchen Chai, Guisheng Fan, Huiqun Yu, Zijie Huang 0001, Jianshu Ding, Yao Guan
Softw. Qual. J.4
2024 Learning to Generate Structured Code Summaries From Hybrid Code Context
abstract
Code summarization aims to automatically generate natural language descriptions for code, and has become a rapidly expanding research area in the past decades. Unfortunately, existing approaches mainly focus on the “one-to-one” mapping from methods to short descriptions, which hinders them from becoming practical tools: 1) The program context is ignored, so they have difficulty in predicting keywords outside the target method; 2) They are typically trained to generate brief function descriptions with only one sentence in length, and therefore have difficulty in providing specific information. These drawbacks are partially due to the limitations of public code summarization datasets. In this paper, we first build a large code summarization dataset including different code contexts and summary content annotations, and then propose a deep learning framework that learns to generate structured code summaries from hybrid program context, named StructCodeSum. It provides both an LLM-based approach and a lightweight approach which are suitable for different scenarios. Given a target method, StructCodeSum predicts its function description, return description, parameter description, and usage description through hybrid code context, and ultimately builds a Javadoc-style code summary. The hybrid code context consists of path context, class context, documentation context and call context of the target method. Extensive experimental results demonstrate: 1) The hybrid context covers more than 70% of the summary tokens in average and significantly boosts the model performance; 2) When generating function descriptions, StructCodeSum outperforms the state-of-the-art approaches by a large margin; 3) According to human evaluation, the quality of the structured summaries generated by our approach is better than the documentation generated by Code Llama.
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001
IEEE Trans. Software Eng.6
2023 Towards Retrieval-Based Neural Code Summarization: A Meta-Learning Approach
abstract
Code summarization aims to generate code summaries automatically, and has attracted a lot of research interest lately. Recent approaches to it commonly adopt neural machine translation techniques, which train a Seq2Seq model on a large corpus and assume it could work on various new code snippets. However, codes are highly varied in practice due to different domains, businesses or programming styles. Therefore, it is challenging to learn such a variety of patterns into a single model. In this paper, we propose a brand-new framework for code summarization based on meta-learning and code retrieval, named MLCS to tackle this issue. In this framework, the summarization of each target code is formalized as a few-shot learning task, where its similar examples are used as training data and the testing example is itself. We retrieve examples similar to the target code in a rank-and-filter manner. Given a neural code summarizer, we optimize it into a meta-learner via Model-Agnostic Meta-Learning (MAML). During inference, the meta-learner first adapts to the retrieved examples and yields an exclusive model for the target code, and then generates its summary. Extensive experiments on real-world datasets show: (1) Utilizing MLCS, a standard Seq2Seq model is able to outperform previous state-of-the-art approaches, including both neural models and retrieval-based neural models; (2) MLCS can flexibly adapt to existing neural code summarizers without modifying their architecture, and could significantly improve their performance with the relative gain of up to 112.7% on BLEU-4, 23.2% on ROUGE-L, and 31.5% on METEOR; (3) Compared to the existing retrieval-based neural approaches, MLCS can better leverage multiple similar examples, and shows better generalization ability on different retrievers, unseen retrieval corpus and low-frequency words.
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004
IEEE Trans. Software Eng.4
2022 Bug Report Priority Prediction Using Developer-Oriented Socio-Technical Features
abstract
Software stakeholders report bugs in Issue Tracking System (ITS) with manually labeled priorities. However, the lack of knowledge and standard for prioritization may cause stakeholders to mislabel the priorities. In response, priority predictors are actively developed to support them. Prior studies trained machine learners based on textual similarity, categorical, and numeric technical features of bug reports. Most models were validated by time-insensitive approaches, and they were producing sub-optimal results for practical usage. Moreover, they tend to ignore the developer and social aspects of ITS. Since ITS bridges users and developers, we integrate their sentiment- and community-oriented socio-technical features to perform 2- and multi-classed bug priority prediction and validate our model in within-project, cross-project, and time-wise scenarios. The proposed model outperforms the 2 baselines by up to 10% in AUC-ROC and 13% in MCC, and the significance of improvement is statistically confirmed. We reveal involving assignee and reporter features from socio-technical perspectives such as sentiment could boost prediction performance. Finally, we test statistically the mean and distribution of the features that reflect the differences in socio-technical aspects (e.g., quality of communication and resource distribution) between high and low priority reports. In conclusion, we suggest researchers should involve contributors’ experience and sentiments in bug report priority prediction.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
Internetware1
2022 HQLgen: deep learning based HQL query generation from program context
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Jiayin Zhang
Autom. Softw. Eng.4
2022 Automatic Identification of High-Impact Bug Report by Product and Test Code Quality
abstract
Bug reports are submitted by the software stakeholders to foster the location and elimination of bugs. However, in large-scale software systems, it may be impossible to track and solve every bug, and thus developers should pay more attention to High-Impact Bugs (HIBs). Previous studies analyzed textual descriptions to automatically identify HIBs, but they ignored the quality of code, which may also indicate the cause of HIBs. To address this issue, we integrate the features reflecting the quality of production (i.e. CK metrics) and test code (i.e. test smells) into our textual similarity based model to identify HIBs. Our model outperforms the compared baseline by up to 39% in terms of AUC-ROC and 64% in terms of F-Measure. Then, we explain the behavior of our model by using SHAP to calculate the importance of each feature, and we apply case studies to empirically demonstrate the relationship between the most important features and HIB. The results show that several test smells (e.g. Assertion Roulette, Conditional Test Logic, Duplicate Assert, Sleepy Test) and product metrics (e.g. NOC, LCC, PF, and ProF) have important contributions to HIB identification.
Jianshu Ding, Guisheng Fan, Huiqun Yu, Zijie Huang 0001
Int. J. Softw. Eng. Knowl. Eng.4
2022 Class Change Prediction by Incorporating Community Smell: An Empirical Study
abstract
To adapt to changing software requirements, developers need to maintain and modify software through code changes. Predicting change-prone code can help developers to reduce the cost of software maintenance in advance. Prior work confirmed code smell intensity is a reliable metric for predicting change-prone classes. Community smell is a derivation of the concept of code smell in open-source software development community, it refers to poor communication and collaboration problems among developers. We add community smell to existing change prediction models, and propose a software class change prediction model integrating process metrics, code smell intensity metrics, anti-pattern metrics and community smell metrics, which takes into account the technicality and organizational aspects of software development. Experimental results demonstrate that when Multilayer Perceptron is used to build a change prediction model, community smell improves the baseline model by 4.4% and 31.5% in terms of [Formula: see text]-Measure and Recall. In addition, community smell improves baseline model performance to a greater extent in terms of Recall and Precision than code smell-related information.
Qingyuan Dou, Zijie Huang 0001
Int. J. Softw. Eng. Knowl. Eng.4
2022 Code Generation with Hybrid of Structural and Semantic Features Retrieval
abstract
Due to the growing need for faster software delivery, code generation has attracted more and more attention, since it could improve code maintainability by providing suggestions for coding. In the model of generating program source code from natural language (NL), the most effective method is to generate an intermediate architecture (such as Abstract Syntax Tree) combined with a deep learning model. However, these models have the following drawbacks: (1) The data structural information is underutilized and the correlation between samples is not considered. (2) Lack of the ability to memorize large and complex structures, so that complex codes cannot be generated correctly. To address these issues, we propose HRCODE model, a code generation architecture based on Hybrid of structural and semantic features Retrieval CODE model. We transform the NL description into an intermediate structure with structural features. Then, the NL and the intermediate structure are embedded into a vector through weight mixing, and we calculate the similarity score between each vector to retrieve the most relevant samples. Finally, the new input is brought into the PLBART model to generate code. Experiments show that HRCODE is at least 4.7% higher than the state-of-the-art models in the ACC metric and at least 10.3% higher in the BLEU-4 score. We have released our code at https://github.com/jesokang/HRCODE.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Ziyi Zhou 0002
Int. J. Softw. Eng. Knowl. Eng.4
2022 Summarizing source code with hierarchical code representation
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Xingguang Yang
Inf. Softw. Technol.4
2022 Community Smell Occurrence Prediction on Multi-Granularity by Developer-Oriented Features and Process Metrics
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Xingguang Yang, Kang Yang 0004
J. Comput. Sci. Technol.1
2022 HBSniff: A static analysis tool for Java Hibernate object-relational mapping code smell detection
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
Sci. Comput. Program.1
2022 A graph sequence neural architecture for code completion with semantic structure features
abstract
Abstract Code completion plays an important role in intelligent software development for accelerating coding efficiency. Recently, the prediction models based on deep learning have achieved good performance in code completion task. However, the existing models cannot avoid three drawbacks: (i) In the existing models, the code representation loses the information (parent–child information between nodes) and lacks many effective features (orientation between nodes). (ii) The known code structure information is not fully utilized, which will cause the model to generate completely irrelevant results. (iii) Simple sequence modeling ignores repeated patterns and structural information. Besides, previous works cannot capture the characteristics of correlation and directionality between nodes. In this paper, we propose a Code Completion approach named CC‐GGNN, which is graph model based on Gated Graph Neural Networks (GGNNs) to address the problems. We introduce a new architecture to obtain the effective code features from code representation. In order to utilize the known information, we propose Classification Mechanism, which classifies the representation of the node using the known parent node and constructs training graph in the model. The experimental results show that our model outperforms the state‐of‐the‐art methods MRR@5 at most 9.2% and ACC at most 11.4% in datasets.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Zijie Huang 0001
J. Softw. Evol. Process.5
2021 An Empirical Study of Model-Agnostic Interpretation Technique for Just-in-Time Software Defect Prediction
Xingguang Yang, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Ziyi Zhou 0002
CollaborateCom (1)4
2021 Predicting Community Smells' Occurrence on Individual Developers by Sentiments
abstract
Community smells appear in sub-optimal software development community structures, causing unforeseen additional project costs, e.g., lower productivity and more technical debt. Previous studies analyzed and predicted community smells in the granularity of community sub-groups using socio-technical factors. However, refactoring such smells requires the effort of developers individually. To eliminate them, supportive measures for every developer should be constructed according to their motifs and working states. Recent work revealed developers' personalities could influence community smells' variation, and their sentiments could impact productivity. Thus, sentiments could be evaluated to predict community smells' occurrence on them. To this aim, this paper builds a developer-oriented and sentiment-aware community smell prediction model considering 3 smells such as Organizational Silo, Lone Wolf, and Bottleneck. Furthermore, it also predicts if a developer quitted the community after being affected by any smell. The proposed model achieves cross- and within-project prediction F-Measure ranging from 76% to 93%. Research also reveals 6 sentimental features having stronger predictive power compared with activeness metrics. Imperative and indicative expressions, politeness, and several emotions are the most powerful predictors. Finally, we test statistically the mean and distribution of sentimental features. Based on our findings, we suggest developers should communicate in a straightforward and polite way.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Ziyi Zhou 0002, Kang Yang 0004, Xingguang Yang
ICPC1
2021 Automatic Identification of High Impact Bug Report by Test Smells of Textual Similar Bug Reports
abstract
Bug reports are written by the software stakeholders to track software defects and vulnerabilities. Since Software Quality Assurance (SQA) resources are limited, developers tend to resolve High-Impact Bugs (HIB) in advance. Prior research identified HIBs by analyzing the textual information in bug reports. However, they only consider textual information instead of the root cause of bugs, such as code quality. Since prior study revealed software test smells (i.e., sub-optimal test code implementation) are related to bug proneness, we intend to measure test smell distribution in textual similar bug reports to identify HIB reports. We first construct an effective model, which outperforms the baseline by 29.3% in terms of AUC-ROC. Secondly, we use SHAP to compute the importance of test smell features. Finally, we conduct an empirical survey to discuss the relationship between test smell and HIB reports. Result shows that Assertion Roulette and Conditional Test Logic test smell are important factors in distinguishing the types of bug reports.
Jianshu Ding, Guisheng Fan, Huiqun Yu, Zijie Huang 0001
QRS4
2021 Python Code Smell Refactoring Route Generation Based on Association Rule and Correlation
abstract
Code smell is a software quality problem caused by software design flaws. Refactoring code smells can improve software maintainability. While prior works mostly focused on Java code smells, only a few prior researches detect and refactor code smells of Python. Therefore, we intend to outline a route (i.e. sequential refactoring operation) for refactoring Python code smells, including LC, LM, LMC, LPL, LSC, LBCL, LLF, MNC, CCC and LTCE. The route could instruct developers to save effort by refactoring the smell strongly correlated with other smells in advance. As a result, more smells could be resolved by a single refactoring. First, we reveal the co-occurrence and the inter-causation between smells. Then, we evaluate the smells’ correlation. Results highlight seven groups of smells with high co-occurrence. Meanwhile, 10 groups of smells correlate with each other in a significant level of Spearman’s correlation coefficient at 0.01. Finally, we generate the refactoring route based on the association rules, we exploit an empirical verification with 10 developers involved. The results of Kendall’s Tau show that the proposed refactoring route has a high inter-agreement with the developer’s perception. In conclusion, we propose four refactoring routes to provide guidance for practitioners, i.e. {LPL [Formula: see text] LLF}, {LPL [Formula: see text] LBCL}, {LPL [Formula: see text] LMC} and {LPL [Formula: see text] LM [Formula: see text] LC [Formula: see text] CCC [Formula: see text] MNC}.
Zijie Huang 0001
Int. J. Softw. Eng. Knowl. Eng.4
2019 The Smell of Blood: Evaluating Anemia and Bloodshot Symptoms in Web Applications
Zijie Huang 0001
SEKE1