VLDB 2026 Research / reviewers in the wild / expert
Yu Wang 0215
dblp:02/5889-215
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-6123-3612ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | iiPCS: Intent-Based In-Context Learning for Project-Specific Code SummarizationabstractRecent years have witnessed growing research interest in automatic source code summarization due to its beneficial potential in software development and maintenance tasks. In the past few years, various deep learning models have been developed to leverage structural and textual features in the code for generating meaningful and succinct summaries. However, the summaries generated by traditional deep learning models often have syntax errors or are meaningless. The emergence of large language models provides an opportunity to overcome the problem. However, the quality of the summaries largely depends on the in-context learning examples of code-summary pairs. In this work, we develop iiPCS, an LLM-based method for code summarization. We retrieve relevant code-summary pairs as in-context learning examples from the same project of the target code, which ensures to generate more project-specific summaries, and use the predicted intent of the target code to pick few-shot examples, which ensures to generate summaries with the correct intent. Experimental results show that iiPCS can generate code summaries with higher quality compared to traditional methods using deep learning and recent methods using LLMs. Yu Wang 0215, Xin Liu 0151, Aoying Zhou |
IJCNN | 1 |
| 2024 | Job Title Prediction as a Dual Task of Expertise Prediction in Open Source Software
Xin Liu 0151, Yu Wang 0215, Qiwen Dong |
ECML/PKDD (10) | 2 |
| 2024 | Code Summarization with Project-Specific Features
Yu Wang 0215, Xin Liu 0151, Aoying Zhou |
ECML/PKDD (9) | 1 |
| 2023 | ErrorCLR: Semantic Error Classification, Localization and Repair for Introductory Programming AssignmentsabstractProgramming education at scale increasingly relies on automated feedback to help students learn to program. An important form of feedback is to point out semantic errors in student programs and provide hints for program repair. Such automated feedback depends essentially on solving the tasks of classification, localization and repair of semantic errors. Although there are datasets for the tasks, we observe that they do not have the annotations supporting all three tasks. As such, existing approaches for semantic error feedback treat error classification, localization and repair as independent tasks, resulting in sub-optimal performance on each task. Moreover, existing datasets either contain few programming assignments or have few programs for each assignment. Therefore, existing approaches often leverage rule-based methods and evaluate them with a small number of programming assignments. To tackle the problems, we first describe the creation of a new dataset COJ2022 that contains 5,914 C programs with semantic errors submitted to 498 different assignments in an introductory programming course, where each program is annotated with the error types and locations and is coupled with the repaired program submitted by the same student. We show the advantages of COJ2022 over existing datasets on various aspects. Second, we treat semantic error classification, localization and repair as dependent tasks, and propose a novel two-stage method ErrorCLR to solve them. Specifically, in the first stage we train a model based on graph matching networks to jointly classify and localize potential semantic errors in student programs, and in the second stage we mask error spans in buggy programs using information of error types and locations and train a CodeT5 model to predict correct spans. The predicted spans replace the error spans to form repaired programs. Experimental results show that ErrorCLR remarkably outperforms the comparative methods for all three tasks on COJ2022 and other public datasets. We also conduct a case study to visualize and interpret what is learned by the graph matching network in ErrorCLR. We have released the source code and COJ2022 at https://github.com/DaSESmartEdu/ErrorCLR. Siqi Han, Yu Wang 0215 |
SIGIR | 2 |
| 2022 | GypSum: learning hybrid representations for code summarizationabstractCode summarization with deep learning has been widely studied in recent years. Current deep learning models for code summarization generally follow the principle in neural machine translation and adopt the encoder-decoder framework, where the encoder learns the semantic representations from source code and the decoder transforms the learnt representations into human-readable text that describes the functionality of code snippets. Despite they achieve the new state-of-the-art performance, we notice that current models often either generate less fluent summaries, or fail to capture the core functionality, since they usually focus on a single type of code representations. As such we propose GypSum, a new deep learning model that learns hybrid representations using graph attention neural networks and a pre-trained programming and natural language model. We introduce particular edges related to the control flow of a code snippet into the abstract syntax tree for graph construction, and design two encoders to learn from the graph and the token sequence of source code, respectively. We modify the encoder-decoder sublayer in the Transformer's decoder to fuse the representations and propose a dual-copy mechanism to facilitate summary generation. Experimental results demonstrate the superior performance of GypSum over existing code summarization models. Yu Wang 0215, Aoying Zhou |
ICPC | 1 |