Deze Wang

dblp:255/1071 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-7935-6840ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 THINK: Tackling API Hallucinations in LLMs via Injecting Knowledge
abstract
Large language models (LLMs) have made significant strides in code generation but often struggle with API hallucination issues, especially for the third-party library. Existing approaches attempt to enhance LLMs by incorporating documentation. However, they face three main challenges: the introduction of irrelevant information that distracts the model; reliance solely on documentation that results in discrepancies between API descriptions and practical usage; and the absence of comprehensive error post-processing mechanisms. To address these challenges, we propose THINK11THINK's benchmark and code is available at https://github.com/Leah-Ljx/think., a knowledge injection method that leverages a custom API knowledge database with two phases: pre-execution enhancement and post-execution optimization. The former reduces irrelevant information and integrates multiple knowledge sources, while the latter identifies seven API error types and suggests three heuristic correction strategies. We manually construct a benchmark by collecting and filtering complex API-related tasks from GitHub to evaluate the effectiveness of our method. The experimental results demonstrate that our method can significantly improve the correctness of API usage in the context of LLMs. We reduce the error rate of programs from 61.18% to 16.64% for GPT-3.5 and from 41.49% to 5.58% for GPT-4o across tasks involving different libraries.
Deze Wang, Yiwei Li 0006, Wei Dong 0006
SANER3
2025 A Little Help Goes a Long Way: Tutoring LLMs in Solving Competitive Programming Through Hints
abstract
Code generation has advanced with large language models (LLMs), but LLMs still struggle with complex tasks, especially in competitive programming. These tasks require understanding complex problems, generating correct code that passes numerous test cases, and meeting tight time and memory limits. We observed that there are some critical hints provided by competition platforms, which often point to the most critical information needed to solve the problem, thus guiding participants to accurate solutions. Inspired by these observations, we propose TEACH1, an approach that tutors LLMs in solving competitive programming by combining critical hints with a structured Chain of Thought (CoT). The key insight of TEACH is to employ a domain-specialized hint generator that is fine-tuned on curated data from competitive programming platforms, enabling it to produce concise and targeted algorithmic hints. By integrating these hints into the reasoning process of LLMs, TEACH helps LLMs bridge the gap between complex tasks and solutions. Furthermore, TEACH simulates human problem-solving through a structured CoT that covers problem understanding, analysis, algorithm selection, and coding. We extensively evaluate TEACH on both proprietary (GPT-3.5, GPT-4o, Claude-3.5-Sonnet, Gemini-2.5-Flash) and open-source (DeepSeek-V3) LLMs. TEACH achieves up to 6.56 absolute (17.4% relative) gain in pass@1 on LeetCode, and demonstrates strong generalization to APPS and ASAC, with maximum pass@1 relative improvements of 17.6% and 26.9%, respectively. Furthermore, existing CoT methods with the hints generated from TEACH yield additional gains, demonstrating its compatibility and extensibility across models and prompting strategies.
Wei Dong 0006, Shangwen Wang, Deze Wang, Tiecheng Ma, Yiwei Li 0006, Kang Yang 0001
IEEE Trans. Software Eng.5
2024 Preference-Guided Refactored Tuning for Retrieval Augmented Code Generation
abstract
Retrieval-augmented code generation utilizes Large Language Models as the generator and significantly expands their code generation capabilities by providing relevant code, documentation, and more via the retriever. The current approach suffers from two primary limitations: 1) information redundancy. The indiscriminate inclusion of redundant information can result in resource wastage and may misguide generators, affecting their effectiveness and efficiency. 2) preference gap. Due to different optimization objectives, the retriever strives to procure code with higher ground truth similarity, yet this effort does not substantially benefit the generator. The retriever and the generator may prefer different golden code, and this gap in preference results in a suboptimal design. Additionally, differences in parameterization knowledge acquired during pre-training result in varying preferences among different generators.
Yun Xiong, Deze Wang, Zhenhan Guan, Zejian Shi, Haofen Wang, Shanshan Li 0001
ASE3
2023 One Adapter for All Programming Languages? Adapter Tuning for Code Search and Summarization
abstract
As pre-trained models automate many code intel-ligence tasks, a widely used paradigm is to fine-tune a model on the task dataset for each programming language. A recent study reported that multilingual fine-tuning benefits a range of tasks and models. However, we find that multilingual fine-tuning leads to performance degradation on recent models UniXcoder and CodeT5. To alleviate the potentially catastrophic forgetting issue in multilingual models, we fix all pre-trained model parameters, insert the parameter-efficient structure adapter, and fine-tune it. Updating only 0.6% of the overall parameters compared to full-model fine-tuning for each programming language, adapter tuning yields consistent improvements on code search and sum-marization tasks, achieving state-of-the-art results. In addition, we experimentally show its effectiveness in cross-lingual and low-resource scenarios. Multilingual fine-tuning with 200 samples per programming language approaches the results fine-tuned with the entire dataset on code summarization. Our experiments on three probing tasks show that adapter tuning significantly outperforms full-model fine-tuning and effectively overcomes catastrophic forgetting.
Deze Wang, Boxing Chen, Shanshan Li 0001, Shaoliang Peng, Wei Dong 0006, Xiangke Liao
ICSE1
2022 Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding
abstract
With the great success of pre-trained models, the pretrain-then-finetune paradigm has been widely adopted on downstream tasks for source code understanding. However, compared to costly training a large-scale model from scratch, how to effectively adapt pre-trained models to a new task has not been fully explored. In this paper, we propose an approach to bridge pre-trained models and code-related tasks. We exploit semantic-preserving transformation to enrich downstream data diversity, and help pre-trained models learn semantic features invariant to these semantically equivalent transformations. Further, we introduce curriculum learning to organize the transformed data in an easy-to-hard manner to fine-tune existing pre-trained models.
Deze Wang, Zhouyang Jia, Shanshan Li 0001, Yue Yu 0001, Yun Xiong, Wei Dong 0006, Xiangke Liao
ICSE1
2021 MulCode: A Multi-task Learning Approach for Source Code Understanding
abstract
Recent years have witnessed the significant rise of Deep Learning (DL) techniques applied to source code. Researchers exploit DL for a multitude of tasks and achieve impressive results. However, most tasks are explored separately, resulting in a lack of generalization of the solutions. In this work, we propose MulCode, a multi-task learning approach for source code understanding that learns unified representation space for tasks, with the pre-trained BERT model for the token sequence and the Tree-LSTM model for abstract syntax trees. Furthermore, we integrate two source code views into a hybrid representation via the attention mechanism and set learnable uncertainty parameters to adjust the tasks' relationship.We train and evaluate MulCode in three downstream tasks: comment classification, author attribution, and duplicate function detection. In all tasks, MulCode outperforms the state-of-the-art techniques. Moreover, experiments on three unseen tasks demonstrate the generalization ability of MulCode compared with state-of-the-art embedding methods.
Deze Wang, Yue Yu 0001, Shanshan Li 0001, Wei Dong 0006, Ji Wang 0001, Qing Liao 0001
SANER1
2020 BugSum: Deep Context Understanding for Bug Report Summarization
abstract
During collaborative software development, bug reports are dynamically maintained and evolved as a part of a software project. For a historical bug report with complicated discussions, an accurate and concise summary can enable stakeholders to reduce the time effort perusing the entire content. Existing studies on bug report summarization, based on whether supervised or unsupervised techniques, are limited due to their lack of consideration of the redundant information and disapproved standpoints among developers' comments. Accordingly, in this paper, we propose a novel unsupervised approach based on deep learning network, called BugSum. Our approach integrates an auto-encoder network for feature extraction with a novel metric (believability) to measure the degree to which a sentence is approved or disapproved within discussions. In addition, a dynamic selection strategy is employed to optimize the comprehensiveness of the auto-generated summary represented by limited words. Extensive experiments show that our approach outperforms 8 comparative approaches over two public datasets. In particular, the probability of adding controversial sentences that are clearly disapproved by other developers during the discussion, into the summary is reduced by up to 69.6%.
Yue Yu 0001, Shanshan Li 0001, Deze Wang, Xiaoguang Mao
ICPC5