VLDB 2026 Research / reviewers in the wild / expert
Yujia Chen 0004
dblp:130/5975-4
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-2901-0643ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Function to Repository: Towards Repository-Level Evaluation of Software Vulnerability DetectionabstractDeep Learning (DL)-based methods have proven to be effective for software vulnerability detection, with a potential for substantial productivity enhancements for detecting vulnerabilities. Current methods mainly focus on detecting single functions (i.e., intra-procedural vulnerabilities), ignoring the more complex inter-procedural vulnerability detection scenarios in practice. For example, developers routinely engage with program analysis to detect vulnerabilities that span multiple functions within repositories. In addition, the widely-used benchmark datasets generally contain only intra-procedural vulnerabilities, leaving the assessment of inter-procedural vulnerability detection capabilities unexplored.To mitigate the issues, we propose a holistic multi-level evaluation system, namedVulEval, aiming at evaluating the detection performance of inter- and intra-procedural vulnerabilities simultaneously. Specifically, VulEval consists of three interconnected evaluation tasks:(1) Function-Level Vulnerability Detection, aiming at detecting intra-procedural vulnerability given a code snippet;(2) Vulnerability-Related Dependency Prediction, aiming at retrieving the vulnerable-related dependency from call graphs for providing developers with explanations about the vulnerabilities; and(3) Repository-Level Vulnerability Detection, aiming at detecting inter-procedural vulnerabilities by combining with the dependencies identified in the second task. VulEval also consists of a large-scale dataset, with a total of 4,196 CVE entries, 232,239 functions, and corresponding 4,699 repository-level source code in C/C++ programming languages. By evaluating 19 vulnerability detection methods on the data split randomly and by time respectively, we observe that the repository-level vulnerability detection framework outperforms the corresponding function-level methods, with an increase of 7.43% in precision, 3.38% in recall, 4.91% in F1 score, and 5.24% in MCC on average except for PILOT. It indicates that incorporating vulnerability-related dependencies facilitates vulnerability detection. Our experimental results also demonstrate that the performance of program-analysis- and prompt-based methods are not affected when splitting the data by time. In addition, our findings indicate that the split setting, retrieval techniques, and vulnerability types have substantial impacts on the performance of repository-level vulnerability detection. We conclude our insights and takeaways for researchers and developers for software vulnerability detection in practice. Xin-Cheng Wen, Xinchen Wang 0001, Yujia Chen 0004, Ruida Hu, David Lo 0001, Cuiyun Gao 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Multi-view Leaderboard: Towards Evaluating the Code Intelligence of LLMs From Multiple ViewsabstractLarge Language Models (LLMs) have shown remarkable performance in code intelligence tasks, prompting the development of various benchmarks and leaderboards to assess their effectiveness across diverse programming scenarios. However, existing leaderboards often rely on coarse-grained metrics and overlook performance variations across different types of tasks. In this paper, we introduce Multi-view Leaderboard, a comprehensive evaluation framework designed to assess the coding capabilities of LLMs from multiple views. Our leaderboard partitions widely-used datasets such as HumanEval, MBPP, and ComplexCodeEval into subsets based on factors like prompt length, problem complexity, and task type. It supports four popular code intelligence tasks including code generation, code completion, test case generation, and API recommendation. Additionally, our leaderboard presents results using ranking tables, line charts, radar charts, and heatmaps. Based on LLMs’ performance on different subsets, we provide model recommendations tailored to different real-world scenarios via a Sankey diagram. A user study involving 11 participants revealed that 90% valued the leaderboard’s practical usefulness for analyzing LLMs’ code intelligence from multiple perspectives. The Multi-view Leaderboard is available at https://huggingface.co/spaces/MVLLL/Multi-view-leaderboard. The demonstration video is available at https://youtu.be/J-zQiOYa1Y8 Zexun Zhan, Cuiyun Gao 0001, Yujia Chen 0004, Guoai Xu, Chun Yong Chong, Shan Gao 0009, Xin Xia 0001 |
APSEC | 4 |
| 2025 | Automated Prompt Generation for Code Intelligence: An Empirical study and Experience in WeChatabstractLarge Code Models (LCMs) have demonstrated potential in advancing various code intelligence tasks. However, their effectiveness can be greatly influenced by the quality of the prompts. Current prompt design strategies in code intelligence studies are mostly manually generated, which could be time-consuming and extremely rely on the base LCMs and tasks. Although automated prompt generation (APG) has been investigated in the natural language processing field, it has not attracted sufficient attention and been well explored in the code intelligence tasks. Considering the various tasks and black-box nature of LCMs faced by developers in practice, it is essential to automate the prompt generation process.To mitigate the gap, we empirically investigate the two important parts in APG, including Instruction Generation (IG) and Muti-Step Reasoning (MSR). The instruction generation part aims at providing a task-related description for instructing LCMs to effectively accomplish specific tasks; while the multistep reasoning part aims at guiding LCMs to produce a series of logical steps before arriving at the final answer. For each part, we evaluate the widely-used APG methods on four open-source LCMs and three code intelligence tasks, i.e., code translation (PL-PL), code summarization (PL-NL) and API recommendation (NL-PL). Experimental results indicate that the two parts in APG can dramatically enhance the performance of the code intelligence tasks compared with the basic prompts. Based on the results, we further propose a novel APG approach by combining the best methods of the two studied parts of APG. Experiments show that the proposed APG approach achieves an average improvement of 28.38% with respect to CodeBLEU for the code translation, 58.11% in terms of ROUGE-L for the code summarization and 84.53% in SuccessRate@1 for the API recommendation over the basic prompts, respectively. To validate the effectiveness in industrial scenario, we further evaluate our approach on WeChat-Bench, a proprietary dataset from the WeChat Group in Tencent for API recommendation, achieving an average improvement of 148.89% in MRR. Kexing Ji, Shiyun Fu, Cuiyun Gao 0001, Yujia Chen 0004, Zezhou Yang, Chaozheng Wang, Yuetang Deng |
ASE | 4 |
| 2024 | Reviewers
Chaozheng Wang, Chunjiong Zhang, Elena Molino-Peña, Jindong Feng, Shuzheng Gao, Xin-Cheng Wen, Yuanchao Liu, Yujia Chen 0004, Zhuofeng Zhao, Zhangbing Zhou, Yucong Duan, Shizhan Chen, Guobing Zou, Buqing Cao |
SSE | 12 |
| 2024 | Bridge and Hint: Extending Pre-trained Language Models for Long-Range CodeabstractIn the field of code intelligence, effectively modeling long-range code poses a significant challenge. Existing pre-trained language models (PLMs) such as UniXcoder have achieved remarkable success, but they still face difficulties with long code inputs. This is mainly due to their limited capacity to maintain contextual continuity and memorize the key information over long-range code. To alleviate the difficulties, we propose EXPO, a framework for EXtending Pre-trained language models for lOng-range code. EXPO incorporates two innovative memory mechanisms we propose in this paper: Bridge Memory and Hint Memory. Bridge Memory uses a tagging mechanism to connect disparate snippets of long-range code, helping the model maintain contextual coherence. Hint Memory focuses on crucial code elements throughout the global context, such as package imports, by integrating a 𝑘NN attention layer to adaptively select the relevant code elements. This dual-memory approach bridges the gap between understanding local code snippets and maintaining global code coherence, thereby enhancing the model’s overall comprehension of long code sequences. We validate the effectiveness of EXPO on five popular pre-trained language models such as UniXcoder and two code intelligence tasks including API recommendation and vulnerability detection. Experimental results demonstrate that EXPO significantly improves the pre-training language models. Yujia Chen 0004, Cuiyun Gao 0001, Zezhou Yang, Hongyu Zhang 0002, Qing Liao 0001 |
ISSTA | 1 |
| 2024 | APIGen: Generative API Method RecommendationabstractAutomatic API method recommendation is an essential task of code intelligence, which aims to suggest suitable APIs for programming queries. Existing approaches can be categorized into two primary groups: retrieval-based and learning-based approaches. Although these approaches have achieved remarkable success, they still come with notable limitations. The retrieval-based approaches rely on the text representation capabilities of embedding models, while the learning-based approaches require extensive task-specific labeled data for training. To mitigate the limitations, we propose APIGen, a generative API recommendation approach through enhanced in-context learning (ICL). APIGen has a powerful representation capability and can make effective recommendations with only a few examples via I CL. To overcome the limitations of standard ICL in capturing task-specific knowledge, APIGen involves two main components: (1) Diverse Examples Selection. APIGen searches for similar posts to the programming queries from the lexical, syntactical, and semantic perspectives, providing more informative examples for ICL. (2) Guided API Recommendation. APIGen enables large language models (LLMs) to perform reasoning before generating API recommendations, where the reasoning involves fine-grained matching between the task intent behind the queries and the factual knowledge of the APIs. With the reasoning process, APIGen makes recommended APIs better meet the programming requirement of queries and also enhances the interpretability of results. We compare APIGen with four existing approaches on two publicly available benchmarks. Experiments show that APIGen outperforms the best baseline CLEAR by 105.8% in method-level API recommendation and 54.3 % in class-level API recommendation in terms of SuccessRate@l. Besides, APIGen achieves an average 49.87 % increase compared to the zero-shot performance of popular LLMs such as GPT-4 in method-level API recommendation regardina the SuccessRate@ 3 metric. Yujia Chen 0004, Cuiyun Gao 0001, Muyijie Zhu, Qing Liao 0001, Yong Wang 0008, Guoai Xu |
SANER | 1 |
| 2023 | API Usage Recommendation Via Multi-View Heterogeneous Graph Representation LearningabstractDevelopers often need to decide which APIs to use for the functions being implemented. With the ever-growing number of APIs and libraries, it becomes increasingly difficult for developers to find appropriate APIs, indicating the necessity of automatic API usage recommendation. Previous studies adopt statistical models or collaborative filtering methods to mine the implicit API usage patterns for recommendation. However, they rely on the occurrence frequencies of APIs for mining usage patterns, thus prone to fail for the low-frequency APIs. Besides, prior studies generally regard the API call interaction graph as homogeneous graph, ignoring the rich information (e.g., edge types) in the structure graph. In this work, we propose a novel method namedMEGAfor improving the recommendation accuracy especially for the low-frequency APIs. Specifically, besidescall interaction graph, MEGA considers another two new heterogeneous graphs:global API co-occurrence graphenriched with the API frequency information andhierarchical structure graphenriched with the project component information. With the three multi-view heterogeneous graphs, MEGA can capture the API usage patterns more accurately. Experiments on three Java benchmark datasets demonstrate that MEGA significantly outperforms the baseline models by at least 19% with respect to the Success Rate@1 metric. Especially, for the low-frequency APIs, MEGA also increases the baselines by at least 55% regarding the Success Rate@1 score. Yujia Chen 0004, Cuiyun Gao 0001, Xiaoxue Ren, Yun Peng 0003, Xin Xia 0001, Michael R. Lyu |
IEEE Trans. Software Eng. | 1 |