Hanwei Qian

dblp:323/7987 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0007-9524-1411ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Prompt Learning for Source Code Summarization
abstract
Source) code summarization is the task of automatically generating natural language summaries (also called comments) for given code snippets. Recently, with the successful application of large language models (LLMs) in numerous fields, software engineering researchers have also attempted to adapt LLMs to solve code summarization tasks. The main adaptation schemes include instruction prompting, taskoriented (full-parameter) fine-tuning, and parameter-efficient fine-tuning (PEFT). However, instruction prompting involves designing crafted prompts and requires users to have professional domain knowledge, while task-oriented fine-tuning requires high training costs, and effective, tailored PEFT methods for code summarization are still lacking. In this paper, we propose an effective prompt learning framework for code summarization called PromptCS. It no longer requires users to rack their brains to design effective prompts. Instead, PromptCS trains a prompt agent that can generate continuous prompts to unleash the potential for LLMs in code summarization. Compared to the human-written discrete prompt, the continuous prompts are produced under the guidance of LLMs and are therefore easier to understand by LLMs. PromptCS is non-invasive to LLMs and freezes the parameters of LLMs when training the prompt agent, which can greatly reduce the requirements for training resources. We evaluate the effectiveness of PromptCS on the CodeSearchNet dataset. Experimental results show that PromptCS significantly outperforms instruction prompting schemes (including zero-shot learning and few-shot learning) on all four widely used metrics, including BLEU, METEOR, ROUGE-L, and SentenceBERT, and is comparable to the task-oriented fine-tuning scheme. In some base LLMs, e.g., CodeGen-Multi-2B and StarCoderBase-1B and -3B, PromptCS even outperforms the task-oriented fine-tuning scheme. More importantly, the training efficiency of PromptCS is faster than the task-oriented fine-tuning scheme, with a more pronounced advantage on larger LLMs. The results of the human evaluation demonstrate that PromptCS can generate more good summaries compared to baselines.
Chunrong Fang, Hanwei Qian, Xia Feng, Weisong Sun
QRS3
2025 Mutual Information Guided Backdoor Mitigation for Pre-Trained Encoders
abstract
Self-supervised learning (SSL) is increasingly attractive for pre-training encoders without requiring labeled data. Downstream tasks built on top of those pre-trained encoders can achieve nearly state-of-the-art performance. The pre-trained encoders by SSL, however, are vulnerable to backdoor attacks as demonstrated by existing studies. Numerous backdoor mitigation techniques are designed for downstream task models. However, their effectiveness is impaired and limited when adapted to pre-trained encoders, due to the lack of label information when pre-training. To address backdoor attacks against pre-trained encoders, in this paper, we innovatively propose a mutual information guided backdoor mitigation technique, named MIMIC(MutualInformation guided backdoorMItigation for pre-trained enCoders). MIMIC uses the potentially backdoored encoder as the teacher network and applies knowledge distillation to create a clean student encoder from it. Different from existing knowledge distillation approaches, MIMIC initializes the student with random weights, inheriting no backdoors from teacher nets. Then MIMIC leverages mutual information between each layer and extracted features to locate where benign knowledge lies in the teacher net, with which distillation is deployed to clone clean features from teacher to student. We craft the distillation loss with two aspects, including clone loss and attention loss, aiming to mitigate backdoors and maintain encoder performance at the same time. Our evaluation conducted on two backdoor attacks in SSL demonstrates that MIMIC can significantly reduce the attack success rate by only utilizing$\leq 5$% of clean pre-training data that is accessible to the defender, surpassing seven state-of-the-art backdoor mitigation techniques. The source code of MIMIC is available athttps://github.com/wssun/MIMIC.
Tingxu Han, Weisong Sun, Chunrong Fang, Hanwei Qian, Zhenyu Chen 0001, Xiangyu Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Exploring Large Language Models for Method Name Prediction
abstract
High-quality method names are crucial for program comprehension and maintenance. However, method naming is a challenging task, especially for inexperienced developers. To alleviate this difficulty, various deep-learning techniques have been proposed to predict appropriate names for given method code bodies. The recent advent of Large Language Models (LLMs) has showcased outstanding performance in Natural Language Processing (NLP), leaving a profound impression on the software engineering (SE) community. LLMs have been applied to several SE tasks, e.g., code summarization, demonstrating their significant potential. However, there has not yet been a systematic and comprehensive investigation into how to adapt LLMs to method name prediction (MNP) tasks and their performance on such tasks. To fill this gap, in this paper, we conduct the first empirical study to understand the capabilities of LLMs in MNP tasks. Firstly, we examine LLMs’ comprehension abilities and interaction patterns in the MNP task, testing various prompts and temperature parameters. Results indicate that LLMs excel under 0 temperature and a few-shot prompt template. Following this, we compare the performance of prompt-guided LLMs and fine-tuned PLMs, with GPT-3.5 performing best among LLMs and CodeT5 leading among PLMs, with minimal performance disparity. SentenceBert with Cosine Similarity (SBCS) has been introduced to quantify the semantic differences between predicted and ground-truth names. Lastly, the manual review highlights human evaluations outperforming F1metrics for LLMs, with issues such as global information lack and dataset quality affecting performance identified through case analysis.
Hanwei Qian, Shaomin Zhu
QRS1
2024 Triangle-oriented Community Detection Considering Node Features and Network Topology
abstract
The joint use of node features and network topology to detect communities is called community detection in attributed networks. Most of the existing work along this line has been carried out through objective function optimization and has proposed numerous approaches. However, they tend to focus only on lower-order details, i.e., capture node features and network topology from node and edge views, and purely seek a higher degree of optimization to guarantee the quality of the found communities, which exacerbates unbalanced communities and free-rider effect. To further clarify and reveal the intrinsic nature of networks, we conduct triangle-oriented community detection considering node features and network topology. Specifically, we first introduce a triangle-based quality metric to preserve higher-order details of node features and network topology, and then formulate so-called two-level constraints to encode lower-order details of node features and network topology. Finally, we develop a local search framework based on optimizing our objective function consisting of the proposed quality metric and two-level constraints to achieve both non-overlapping and overlapping community detection in attributed networks. Extensive experiments demonstrate the effectiveness and efficiency of our framework and its potential in alleviating unbalanced communities and free-rider effect.
Guangliang Gao, Weichao Liang, Hanwei Qian, Jie Cao 0001
ACM Trans. Web4
2023 Abstract Syntax Tree for Method Name Prediction: How Far Are We?
abstract
Method name prediction (MNP) aims to recommend a proper name for a method given by the developer, which can ease the programming task and improve programmer productivity. Due to the excellent expressiveness of code representation, abstract syntax trees (AST) have been widely exploited by MNP techniques. However, it is a complex process to manipulate AST, including AST parsing, AST preprocessing, and AST encoding, of which a change in the scheme may change the AST embeddings and thus affects the performance of MNP. In this paper, we first conduct a comprehensive empirical study to systematically investigate the impact of the sub-processes of AST usage on MNP performance. The empirical findings of this study unmistakably demonstrate that AST has a positive impact on promoting MNP. Moreover, the selection of schemes for AST parsing, preprocessing, and encoding exerts a profound influence on the effectiveness of MNP. Properly combining AST (e.g., using JDT, AST Pathfull, and code2seq as AST parsing, preprocessing, and encoding methods, respectively) can improve MNP performance by 164% in terms of F1-score compared to using only code tokens.
Hanwei Qian, Weisong Sun, Chunrong Fang
QRS1