VLDB 2026 Research / reviewers in the wild / expert
Zejian Shi
dblp:202/4905
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0000-1312-7738ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enhancing Code Generation Through Retrieval of Cross-Lingual Semantic GraphsabstractIn the field of software engineering automation, code language models have made significant strides in code generation tasks. However, due to the cost of updating knowledge and the issue of hallucinations, code language models (CLMs) face challenges in practical code generation scenarios, making retrieval-augmented code generation a mainstream approach. Existing retrieval-augmented methods only build codebases for a single programming language, which is insufficient to address the lack of monolingual knowledge. To address this, we propose CodeRCSG, a novel cross-lingual retrieval-augmented code generation method. This method constructs a multilingual codebase and creates a unified cross-lingual code semantic graph to capture deep semantic information across different programming languages. By encoding the retrieved code semantic graph with GNN and combining it with input text embeddings, code language models can effectively utilize the transferred cross-lingual programming knowledge to improve the quality of generated code. Experimental results show that CodeRCSG can significantly enhance the code generation capabilities of code language models. Zhijie Jiang, Zejian Shi, Yun Xiong |
APSEC | 2 |
| 2024 | Preference-Guided Refactored Tuning for Retrieval Augmented Code GenerationabstractRetrieval-augmented code generation utilizes Large Language Models as the generator and significantly expands their code generation capabilities by providing relevant code, documentation, and more via the retriever. The current approach suffers from two primary limitations: 1) information redundancy. The indiscriminate inclusion of redundant information can result in resource wastage and may misguide generators, affecting their effectiveness and efficiency. 2) preference gap. Due to different optimization objectives, the retriever strives to procure code with higher ground truth similarity, yet this effort does not substantially benefit the generator. The retriever and the generator may prefer different golden code, and this gap in preference results in a suboptimal design. Additionally, differences in parameterization knowledge acquired during pre-training result in varying preferences among different generators. Yun Xiong, Deze Wang, Zhenhan Guan, Zejian Shi, Haofen Wang, Shanshan Li 0001 |
ASE | 5 |
| 2023 | Improving Code Search with Multi-Modal Momentum Contrastive LearningabstractContrastive learning has recently been applied to enhancing the BERT-based pre-trained models for code search. However, the existing end-to-end training mechanism cannot sufficiently utilize the pre-trained models due to the limitations on the number and variety of negative samples. In this paper, we propose MoCoCS, a multi-modal momentum contrastive learning method for code search, to improve the representations of query and code by constructing large-scale multi-modal negative samples. MoCoCS increases the number and the variety of negative samples through two optimizations: integrating multi-batch negative samples and constructing multi-modal negative samples. We first build momentum contrasts for query and code, which enables the construction of large-scale negative samples out of a mini-batch. Then, to incorporate multi-modal code information, we build multi-modal momentum contrasts by encoding the abstract syntax tree and the data flow graph with a momentum encoder. Experiments on CodeSearchNet with six programming languages demonstrate that our method can further improve the effectiveness of pre-trained models for code search. Zejian Shi, Yun Xiong, Yao Zhang 0009, Zhijie Jiang, Jinjing Zhao, Shanshan Li 0001 |
ICPC | 1 |
| 2022 | Cross-Modal Contrastive Learning for Code SearchabstractCode search aims to retrieve code snippets from natural language queries, which serves as a core technology to improve development efficiency. Previous approaches have achieved promising results to learn code and query representations by using BERT-based pre-trained models which, however, leads to semantic collapse problems, i.e. native representations of code and query clustering in a high similarity interval. In this paper, we propose CrossCS, a cross-modal contrastive learning method for code search, to improve the representations of code and query by explicit fine-grained contrastive objectives. Specifically, we design a novel and effective contrastive objective that considers not only the similarity between modalities, but also the similarity within modalities. To maintain semantic consistency of code snippets with different names of functions and variables, we use data augmentation to rename functions and variables to meaningless tokens, which enables us to add comparisons between code and augmented code within modalities. Moreover, in order to further improve the effectiveness of pre-trained models, we rank candidate code snippets using similarity scores weighted by retrieval scores and classification scores. Comprehensive experiments demonstrate that our method can significantly improve the effectiveness of pre-trained models for code search. Zejian Shi, Yun Xiong, Yao Zhang 0009, Shanshan Li 0001, Yangyong Zhu |
ICSME | 1 |
| 2017 | The prediction of character based on recurrent neural network language modelabstractThis paper mainly talks about the Recurrent Neural Network and introduces a more effective neural network model named LSTM. Then, the paper recommends a special language model based on Recurrent Neural Network. With the help of LSTM and RNN language models, program can predict the next character after a certain character. The main purpose of this paper is to compare the LSTM model with the standard RNN model and see their results in character prediction. So we can see the huge potential of Recurrent Neural Network Language Model in the field of character prediction. Zejian Shi, Minyong Shi, Chunfang Li |
ICIS | 1 |