VLDB 2026 Research / reviewers in the wild / expert
Fengji Zhang
dblp:287/8086
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-0965-417XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | R2ComSync: improving code-comment synchronization with in-context learning and reranking
Zhen Yang 0022, Xiao Yu 0008, Jacky W. Keung, Shuo Liu 0020, Pak Yuen Patrick Chan, Yicheng Sun, Fengji Zhang |
Empir. Softw. Eng. | 8 |
| 2025 | Towards Engineering Multi-Agent LLMs: A Protocol-Driven ApproachabstractThe increasing demand for software development has driven interest in automating software engineering (SE) tasks using Large Language Models (LLMs). Recent efforts extend LLMs into multi-agent systems (MAS) that emulate collaborative development workflows, but these systems often fail due to three core deficiencies: under-specification, coordination misalignment, and inappropriate verification, arising from the absence of foundational SE structuring principles. This paper introduces Software Engineering Multi-Agent Protocol (SEMAP), a protocol-layer methodology that instantiates three core SE design principles for multi-agent LLMs: (1) explicit behavioral contract modeling, (2) structured messaging, and (3) lifecycleguided execution with verification, and is implemented atop Google’s Agent-to-Agent (A2A) infrastructure. Empirical evaluation using the Multi-Agent System Failure Taxonomy (MAST) framework demonstrates that SEMAP effectively reduces failures across different SE tasks. In code development, it achieves up to a $69.6 \%$ reduction in total failures for function-level development and $\mathbf{5 6 . 7 \%}$ for deployment-level development. For vulnerability detection, SEMAP reduces failure counts by up to $47.4 \%$ on Python tasks and $28.2 \%$ on $\mathrm{C} / \mathrm{C}++$ tasks. Zhenyu Mao, Jacky W. Keung, Fengji Zhang, Shuo Liu 0020 |
APSEC | 3 |
| 2025 | Exploring continual learning in code intelligence with domain-wise distilled prompts
Shuo Liu 0020, Jacky W. Keung, Zhen Yang 0022, Fang Liu 0032, Fengji Zhang, Yicheng Sun |
Inf. Softw. Technol. | 5 |
| 2025 | Private-library-oriented code generation with large language models
Daoguang Zan, Bei Chen 0008, Yongshun Gong, Junzhi Cao, Fengji Zhang, Bingchao Wu, Bei Guan, Yilong Yin, Yongji Wang 0002 |
Knowl. Based Syst. | 5 |
| 2024 | Data preparation for Deep Learning based Code Smell Detection: A systematic literature review
Fengji Zhang, Zexian Zhang, Jacky W. Keung, Xiangru Tang, Zhen Yang 0022, Xiao Yu 0008 |
J. Syst. Softw. | 1 |
| 2023 | Large Language Models Meet NL2Code: A SurveyabstractDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, Jian-Guang Lou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Daoguang Zan, Bei Chen 0008, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang 0002, Jian-Guang Lou |
ACL (1) | 3 |
| 2023 | RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationabstractFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, Weizhu Chen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Fengji Zhang, Bei Chen 0008, Jacky W. Keung, Jin Liu 0016, Daoguang Zan, Jian-Guang Lou, Weizhu Chen |
EMNLP | 1 |
| 2023 | CodeT: Code Generation with Generated Tests
Bei Chen 0008, Fengji Zhang, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, Weizhu Chen |
ICLR | 2 |
| 2023 | Uncovering and Quantifying Social Biases in Code GenerationabstractWith the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the social bias problem in pre-trained code generation models. We propose a new paradigm to construct code prompts and successfully uncover social biases in code generation models. To quantify the severity of social biases in generated code, we develop a dataset along with three metrics to evaluate the overall social bias and fine-grained unfairness across different demographics. Experimental results on three pre-trained code generation models (Codex, InCoder, and CodeGen) with varying sizes, reveal severe social biases. Moreover, we conduct analysis to provide useful insights for further choice of code generation models with low social bias. Yan Liu 0002, Xiaokang Chen, Yan Gao 0002, Fengji Zhang, Daoguang Zan, Jian-Guang Lou, Tsung-Yi Ho |
NeurIPS | 5 |
| 2023 | Diverse title generation for Stack Overflow posts with multiple-sampling-enhanced transformer
Fengji Zhang, Jin Liu 0016, Yao Wan 0001, Xiao Yu 0008, Xiao Liu 0004, Jacky W. Keung |
J. Syst. Softw. | 1 |
| 2022 | Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information
Fengji Zhang, Xiao Yu 0008, Jacky W. Keung, Zhiwen Xie, Zhen Yang 0022, Caoyuan Ma, Zhimin Zhang 0008 |
Inf. Softw. Technol. | 1 |