VLDB 2026 Research / reviewers in the wild / expert
Wentao Zou
dblp:196/8781
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Retrieval-Augmented Code Generation of Low-Resource Programming LanguagesabstractThe performance of Large Language Models degrades substantially when generating code for low-resource programming languages. While Retrieval-Augmented Generation (RAG) offers a solution, applying it to this domain presents unique challenges in knowledge retrieval and integration. To address this, we introduce PEARL, a novel framework for lowresource languages. PEARL constructs high-quality knowledge databases and employs a distillation method to train a retriever using the LLM’s own preferences, bypassing the need for manual annotation. By effectively integrating this external knowledge, PEARL improves performance of LLMs on lowresource programming languages. In evaluations across five low-resource languages, PEARL outperformed RAG baselines, increasing average Pass@1 by 22% on LLaMA-3.1-8B-Instruct and 10% on DeepSeek-Coder-6.7B-Instruct. Jianbo Lin, Chuanyi Li, Wentao Zou, Jidong Ge, Bin Luo 0003 |
APSEC | 4 |
| 2025 | Improving Source Code Pre-Training via Type-Specific MaskingabstractThe Masked Language Modeling (MLM) task is widely recognized as one of the most effective pre-training tasks and currently derives many variants in the Software Engineering (SE) field. However, most of these variants mainly focus on code representation without distinguishing between different code token types, while some focus on a specific type, such as code identifiers. Indeed, various code token types exist, and there is no evidence that only identifiers can improve PTMs. Thus, to improve PTMs through different types, we conducted an extensive study to evaluate how different type-specific masking tasks can affect PTMs. First, we extract five code token types, convert them into type-specific masking tasks, and generate their combinations. Second, we pre-train CodeBERT and PLBART using combinations and fine-tuned them on four SE downstream tasks. Experimental results show that type-specific masking tasks can enhance CodeBERT and PLBART on all downstream tasks. Furthermore, we discuss topics related to low-resource datasets, conflicting PTMs that original pre-training tasks conflict with our methods, the cost and performance of our methods, factors that impact the performance of our methods, and applying our methods on state-of-the-art PTMs. These discussions comprehensively analyze the strengths and weaknesses of different type-specific masking tasks. Wentao Zou, Chuanyi Li, Jidong Ge, Xiang Chen 0005, LiGuo Huang, Bin Luo 0003 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | Experimental Evaluation of Parameter-Efficient Fine-Tuning for Software Engineering TasksabstractPre-trained models (PTMs) have succeeded in various software engineering (SE) tasks following the “pre-train then fine-tune” paradigm. As fully fine-tuning all parameters of PTMs can be computationally expensive, a potential solution is parameter-efficient fine-tuning (PEFT), which freezes PTMs while introducing extra parameters. Although PEFT methods have been applied to SE tasks, researchers often focus on specific scenarios and lack a comprehensive comparison of PTMs from different aspects such as field, size, and architecture. To fill this gap, we have conducted an empirical study on six PEFT methods, eight PTMs, and four SE tasks. The experimental results reveal several noteworthy findings. For example, model architecture has little impact on PTM performance when using PEFT methods. Additionally, we provide a comprehensive discussion of PEFT methods from three perspectives. First, we analyze the effectiveness and efficiency of PEFT methods. Second, we explore the impact of the scaling factor hyperparameter. Finally, we investigate the application of PEFT methods on the latest open source large language model, Llama 3.2. These findings provide valuable insights to guide future researchers in effectively applying PEFT methods to SE tasks. Wentao Zou, Zongwen Shen, Jidong Ge, Chuanyi Li, Xiang Chen 0005, Xiaoyu Shen 0001, LiGuo Huang, Bin Luo 0003 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | CCAF: Learning Code Change via AdapterFusionabstractCode changes are crucial because all code repositories can be viewed as composed of a series of code changes. Recent works on code changes prefer to use pre-trained models (PTMs) to capture the code change representations and have achieved remarkable success. However, these works usually compromise the original code representations of PTMs and ignore the relation of different code-change-related tasks. To boost the existing solutions to code-change-related tasks, we propose a new two-stage Code Change representation learning method using AdapterFusion, which is called CCAF. The first stage is knowledge extraction, where we freeze the parameters of the PTM and fine-tune additional parameters known as adapters. Each adapter acquires knowledge from a specific code-change-related task. The second stage, knowledge composition, employs AdapterFusion to compose the knowledge from all adapters, enhancing the PTM’s performance on a specific code-change-related task. To assess the effectiveness of CCAF, we employ CodeT5 as the base PTM, with its parameters frozen, and apply CCAF to three code-change-related tasks: commit message generation, automated patch correctness assessment, and just-in-time defect prediction. The experimental results indicate that CCAF not only outperforms a fully fine-tuned CodeT5 but also performs comparably to the state-of-the-art method, CCRep. Wentao Zou, Zongwen Shen, Jidong Ge, Chuanyi Li, Bin Luo 0003 |
Internetware | 1 |
| 2024 | Eyeglass Reflection Removal With Joint Learning of Reflection Elimination and Content InpaintingabstractEyeglass reflection removal is of great importance to the portrait image processing. However, it remains a challenge to eliminate the reflections on the glass and restore the textual contents of eyes without introducing visual artifacts. Addressing this problem, in this paper, we propose an Eyeglass Reflection Removal Network (ER2Net) by learning reflection elimination and content inpainting jointly. The reflection elimination branch is effective in weak reflection regions, and the content inpainting branch is dedicated to content reasoning in strong reflection regions. We then propose a result fusion module (RFM), which adaptively fuses the elimination result and the inpainting result according to the reflection intensity of each pixel, to produce high-quality result. We also design a memory module for improving the content inpainting result, and propose an eye-symmetry loss to avoid visual artifacts. Additionally, we construct the first Real-world eyeglass Reflection (ReyeR) dataset for eyeglass reflection removal. Extensive quantitative and qualitative experiments demonstrate the superiority of the ER2Net over state-of-the-art methods for eyeglass reflection removal. Wentao Zou, Xiao Lu 0002, Zhilv Yi, Ling Zhang 0017, Gang Fu 0003, Ping Li 0016, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Root Cause Analysis for Microservice Systems via Hierarchical Reinforcement Learning from Human FeedbackabstractIn microservice systems, the identification of root causes of anomalies is imperative for service reliability and business impact. This process is typically divided into two phases: (i)constructing a service dependency graph that outlines the sequence and structure of system components that are invoked, and (ii) localizing the root cause components using the graph, traces, logs, and Key Performance Indicators (KPIs) such as latency. However, both phases are not straightforward due to the highly dynamic and complex nature of the system, particularly in large-scale commercial architectures like Microsoft Exchange. Lu Wang 0029, Chaoyun Zhang, Ruomeng Ding, Yong Xu 0010, Wentao Zou, Qingjun Chen, Meng Zhang 0025, Xuedong Gao, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
KDD | 6 |
| 2018 | An Image Rain Removal algorithm based on the depth of field and sparse codingabstractRainfall weather can always seriously deteriorate the quality of the outdoor monitoring system image. Since the decomposition based methods do not need to impose any restrictions on the types of rain, they have a wider application in removing the rain streaks. However, they still have the problems of rain residues in the low frequency component, and mis-matching the background and the rain streaks with the same gradient in the high frequency. In this condition, we propose an image rain removal algorithm based on the depth of field and sparse coding. The algorithm includes four steps: image decomposition, dictionary learning, atomic clustering based on Principal Component Analysis and Support Vector Machine, image revising based on the depth of field saliency map. Firstly, the image is decomposed by using the combination of bilateral filtering and short-time Fourier transform, so that the contour in the low-frequency part of the image can be better preserved. The depth of field saliency map of the image is utilized to eliminate the rain residues in the low frequency components, and also to solve the problem of mis-matching the background and the rain streaks with the same gradient in the high frequency components. The experimental results demonstrate that the proposed algorithm performs better both in rain removal and preserving the detailed information of the image than current methods. Junfeng Lei, Shangyue Zhang, Wentao Zou, Jinsheng Xiao, Yunhua Chen, Haigang Sui |
ICPR | 3 |
| 2018 | Video denoising algorithm based on improved dual-domain filtering and 3D block matchingabstractThis study introduces an algorithm for video denoising based on improved dual‐domain filtering and 3D block matching. The wavelet thresholding based on 3D block matching is introduced to make full use of the correlation of video sequence in order to apply dual‐domain filtering to the video. A layered approach is used that attempts denoising in both a base layer and a detail layer. The result of wavelet thresholding based on 3D block matching is used as a guide image to make the base layer smoother. Shrinkage of short‐time Fourier transform coefficients further decreases the noise in the detail layer. Experimental results show that the authors’ algorithm generates a better base layer and detail layer than the traditional dual‐domain filtering algorithms. The subjective and objective comparisons of different algorithms also prove that the proposed algorithm performs better for video denoising. Jinsheng Xiao, Wentao Zou, Shangyue Zhang, Junfeng Lei, Yuan-Fang Wang |
IET Image Process. | 2 |
| 2018 | Single image rain removal based on depth of field and sparse coding
Jinsheng Xiao, Wentao Zou, Yunhua Chen, Junfeng Lei |
Pattern Recognit. Lett. | 2 |
| 2017 | Lane Detection Based on Road Module and Extended Kalman Filter
Jinsheng Xiao, Wentao Zou, Reinhard Klette |
PSIVT | 4 |