Liuwen Cao

dblp:309/6441 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0001-9680-4016ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SRACG: A Code Generation Framework with Selective Retrieval Augmentation
abstract
Large Language Models (LLMs) have demonstrated remarkable performance in code generation, offering new possibilities for translating natural language into executable programs. To further enhance LLMs’ code generation capabilities, Retrieval-Augmented Generation (RAG) has emerged as a promising strategy by retrieving code examples aligned with the generation intent to guide the process. However, existing RAG-based methods often suffer from unnecessary augmentation, preference misalignment, and surface-level mimicry, which undermine the effectiveness of retrieved examples in guiding LLMs toward accurate code generation. To address these challenges, we propose SRACG, a Selective Retrieval-Augmented Code Generation framework. SRACG begins with a necessity-aware selection mechanism to identify generation intents that genuinely require retrieval support, thereby avoiding degradation from indiscriminate augmentation. For intents identified as needing enhancement, it first employs a multi-objective retrieval strategy to select examples that are semantically aligned with the intent. These candidates are then further filtered by assessing their consistency with the LLM’s inherent generation preferences, ensuring alignment in both style and structure. Finally, it extracts execution plans from the filtered examples to uncover their underlying logic, guiding the LLM to better comprehend the examples instead of merely mimicking surface-level content. Experimental results on widely used benchmarks show that SRACG significantly improves the success rate of LLM-generated code and outperforms existing approaches.
Mengzhen Wang, Shukai Ma, Songwen Gong, Jiexin Wang 0002, Ruolin Chen, Liuwen Cao, Yi Cai 0001
AAAI6
2026 A Code Ranking Framework with Human-Inspired Agent-Based Tests Generation
abstract
Pre-trained large language models (LLMs) have emerged as a breakthrough technology in code intelligence such as code generation. Recently, many works have found that LLMs can generate a correct code solution when it is allowed to make numerous attempts. Consequently, a recent trend is to do a large-scale sampling of codes from LLMs and then rank the code to select the most suitable code, a process called code ranking. A common code ranking approach ranks the code by running it against a set of LLM-generated test cases in the form of assert statements. However, existing approaches overlook how humans design test cases through systematic behaviors, which are essential for creating reliable tests for code ranking. Moreover, humans often use input–output examples to clarify and articulate the intended functionality, rather than merely to rank code. To address these gaps, we propose RankAgent, a human-inspired agent-based framework that systematically simulates human test design behaviors. Extensive experiments on five LLMs (including both open- and closed-source models) and two benchmarks (HumanEval+ and LiveCodeBench) show that RankAgent achieves notable and consistent improvements.
Liuwen Cao, Jiexin Wang 0002, Yi Cai 0001
ICPC1
2025 Rethinking-based Code Summarization with Chain of Comments
abstract
Automatic code summarization aims to generate concise natural language descriptions (summary) for source code, which can free software developers from the heavy burden of manual commenting and software maintenance. Existing methods focus on learning a direct mapping from pure code to summaries, overlooking the significant heterogeneity gap between code and summary. Moreover, existing methods lack a human-like re-check process to evaluate whether the generated summaries match well with the code. To address these two limitations, we introduce RBCoSum, a novel framework that incorporates the generated Chain Of Comments (COC) as auxiliary intermediate information for the model to bridge the gap between code and summaries. Also, we propose a rethinking process where a learned ranker trained on our constructed ranking dataset scores the extent of matching between the generated summary and the code, selecting the highest-scoring summary to achieve a re-check process. We conduct extensive experiments to evaluate our approach and compare it with other automatic code summarization models as well as multiple code Large Language Models (LLMs). The experimental results show that RBCoSum is effective and outperforms baselines by a large margin. The human evaluation also proves the summaries generated with RBCoSum are more natural, informative, useful, and truthful.
Liuwen Cao, Hongkui He, Hailin Huang, Jiexin Wang 0002, Yi Cai 0001
COLING1
2025 CLCoSum: Curriculum Learning-Based Code Summarization for Code Language Models
abstract
The code summarization task aims to automatically generate natural language descriptions for code snippets. Recently, pre-trained code language models (CLMs) have demonstrated outstanding performance on code summarization. Additionally, researchers have shown that there is a strong correlation between code function names and summaries, and poorly defined function names lead to worse summaries generated by models. To mitigate this issue, in this paper, we propose CLCoSum, a curriculum learning-based code summarization method for CLMs that improves their performance in poorly named function scenarios. CLCoSum helps CLMs avoid overreliance on function names when they are poorly defined. First, CLCoSum employs data augmentation operators on function names to generate semantically equivalent poorly named codes, which are considered harder data and assist in reducing the model's reliance on unclear function names. Subsequently, CLCoSum uses a curriculum learning paradigm to allow the model to learn these harder codes in an organized way during finetuning. This approach enables CLMs to progress from easier to more difficult training data, similar to the human learning process. Extensive experiments on two existing datasets for Java and Python demonstrate that CLCoSum boosts the performance of various CLMs in code summarization. Specifically, the improvements in BLEU-4 score range from approximately 5% to$\mathbf{2 0. 8 \%}$. The fine-tuning speed of CLCoSum on the augmented dataset is also competitive. Our code and data are available at https://github.com/KuiH/CLCoSum.
Hongkui He, Jiexin Wang 0002, Liuwen Cao, Yi Cai 0001
ICPC3
2025 Code Ranking with Structure Awareness Contrastive Learning
abstract
Large language models (LLMs) have revolutionized the field of programming for developers by automatically generating code based on natural language intent (NL intent). In numerous cases, LLMs can produce correct programs after several trials. As a result, a major challenge for this task is to select the most appropriate program from the multiple samples (also called code ranking) generated by LLMs. Recent popular approaches for code ranking involve the ranker-based methods, in which we train a ranker to classify the error in code using execution results (correct or error types) of code as supervised signals select the best program. However, existing rankerbased code ranking approaches rely on classification labels, which are highly sensitive to label distribution and show weak generalization ability to other distributions. In this paper, we introduce SACL-CR to address this challenge, a novel structureaware contrastive learning framework for code ranking. This approach effectively addresses the generalization issues of existing ranker-based methods by integrating both code sequence and structural information. Encoders trained with this method can effectively identify errors in code, enhancing the model's ability to differentiate between correct and incorrect code. Our research demonstrates that SACL-CR significantly enhances the pass@k accuracy of several code generation models, including CodeLlama and DeepseekCoder, on the HumanEval and MBPP datasets. The open-source code will be released at https://github.com/Iced-Americano2001/SACL-CR.
Hailin Huang, Liuwen Cao, Jiexin Wang 0002, Tianchen Yu, Yi Cai 0001
ICPC2
2024 Beyond Code: Evaluate Thought Steps for Complex Code Generation
abstract
Code generation aims to generate code in a general-purpose programming language, such as C++, based on natural language intents. Existing efforts primarily focus on relatively simple programming problems and fail to evaluate the thought process involved in complex programming scenarios. In this paper, we introduce “steps-guided code generation,” a task that assesses the quality of both thought steps and code implementation to evaluate the overall management of handling a complex programming problem. To support this task, we construct CodeStepsEval, a real-world scenario dataset of complex programming problems in the C++ programming language with varying levels of difficulty. Comprehensive experiments on this dataset demonstrate the importance of high-quality steps in enhancing code generation performance and the challenges faced by the code LLMs in this task.
Liuwen Cao, Yi Cai 0001, Jiexin Wang 0002, Hongkui He, Hailin Huang
LREC/COLING1
2024 Modeling different effects of user and product attributes on review sentiment classification
Changxing Wu, Liuwen Cao, Yuanyun Wang, Jinsong Su
Appl. Intell.2
2022 A Label Dependence-Aware Sequence Generation Model for Multi-Level Implicit Discourse Relation Recognition
abstract
Implicit discourse relation recognition (IDRR) is a challenging but crucial task in discourse analysis. Most existing methods train multiple models to predict multi-level labels independently, while ignoring the dependence between hierarchically structured labels. In this paper, we consider multi-level IDRR as a conditional label sequence generation task and propose a Label Dependence-aware Sequence Generation Model (LDSGM) for it. Specifically, we first design a label attentive encoder to learn the global representation of an input instance and its level-specific contexts, where the label dependence is integrated to obtain better label embeddings. Then, we employ a label sequence decoder to output the predicted labels in a top-down manner, where the predicted higher-level labels are directly used to guide the label prediction at the current level. We further develop a mutual learning enhanced training method to exploit the label dependence in a bottom-up direction, which is captured by an auxiliary decoder introduced during training. Experimental results on the PDTB dataset show that our model achieves the state-of-the-art performance on multi-level IDRR. We release our code at https://github.com/nlpersECJTU/LDSGM.
Changxing Wu, Liuwen Cao, Yubin Ge, Yang Liu 0005, Min Zhang 0005, Jinsong Su
AAAI2