VLDB 2026 Research / reviewers in the wild / expert
Tianchen Yu
dblp:407/8681
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multiple Representation Transformer with Optimized Abstract Syntax Tree for Efficient Code Clone DetectionabstractOver the past decade, the application of deep learning in code clone detection has produced remarkable results. However, the current approaches have two limitations: (a) code representation approaches with low information utilization, such as vanilla Abstract Syntax Tree (AST), leading to information redundancy which results in performance degradation; (b) low efficiency of clone detection on evaluation, resulting in excessive time costs during practical use. In this paper, we propose a Multiple Representation Transformer with an Optimized Abstract Syntax Tree (MRT-OAST) to introduce an efficient code representation method while achieving competitive performance. Specifically, MRT-OAST strategically prunes and enhances the AST, utilizing both pre-order and post-order traversals to represent two different representations. To speed up the evaluation process, MRT-OAST utilizes a pure Siamese Network and employs cosine similarity to compare the similarity between codes. Our approach effectively reduces AST sequences to 40 % and 39 % of their original length in Java and C/C++ while preserving structural information. In code clone detection tasks, our model surpasses state-of-the-art approaches on OJClone and Google Code Jam. During the evaluation of BigCloneBench, our model has a 5x speed improvement compared to the state-of-the-art lightweight model and a 563x speed improvement compared to the BERT-based model, with only a 0.3 % and 0.9 % decrease in$F_{1}$-score. Tianchen Yu, Liannan Lin, Hongkui He |
ICSE | 1 |
| 2025 | Code Ranking with Structure Awareness Contrastive LearningabstractLarge language models (LLMs) have revolutionized the field of programming for developers by automatically generating code based on natural language intent (NL intent). In numerous cases, LLMs can produce correct programs after several trials. As a result, a major challenge for this task is to select the most appropriate program from the multiple samples (also called code ranking) generated by LLMs. Recent popular approaches for code ranking involve the ranker-based methods, in which we train a ranker to classify the error in code using execution results (correct or error types) of code as supervised signals select the best program. However, existing rankerbased code ranking approaches rely on classification labels, which are highly sensitive to label distribution and show weak generalization ability to other distributions. In this paper, we introduce SACL-CR to address this challenge, a novel structureaware contrastive learning framework for code ranking. This approach effectively addresses the generalization issues of existing ranker-based methods by integrating both code sequence and structural information. Encoders trained with this method can effectively identify errors in code, enhancing the model's ability to differentiate between correct and incorrect code. Our research demonstrates that SACL-CR significantly enhances the pass@k accuracy of several code generation models, including CodeLlama and DeepseekCoder, on the HumanEval and MBPP datasets. The open-source code will be released at https://github.com/Iced-Americano2001/SACL-CR. Hailin Huang, Liuwen Cao, Jiexin Wang 0002, Tianchen Yu, Yi Cai 0001 |
ICPC | 4 |
| 2025 | Mixture-of-Experts Low-Rank Adaptation for Multilingual Code SummarizationabstractAs Code Language Models (CLMs) are increasingly used to automate multilingual code intelligence tasks, Full-Parameter Fine-Tuning (FPFT) of CLMs has become a widely adopted approach, which is both time-consuming and resource-intensive. Parameter-Efficient Fine-Tuning (PEFT) provides a more efficient alternative to FPFT. However, it struggles to capture common features shared across languages, leading to performance degradation. Recent studies have explored mixed-language training with PEFT to avoid the loss of common features. However, these methods can result in gradient conflicts due to the diverse language-specific features, causing suboptimal performance, particularly for low-resource languages. In this paper, we propose Mixture-of-Experts Multilingual Low-Rank Adaptation (MMLoRA) for multilingual code summarization. MMLoRA addresses gradient conflicts while preserving common features shared across languages by combining a universal expert with a set of specialized linguistic experts. Additionally, we introduce an expert loss function that maintains the diversity of specialized linguistic experts while balancing the learning progress. Experimental results indicate that MMLoRA achieves state-of-the-art performance in multilingual code summarization while maintaining efficient fine-tuning. The performance improvement is particularly significant in low-resource languages such as Ruby. Tianchen Yu, Hailing Huang, Jiexin Wang 0002, Yi Cai 0001 |
ASE | 1 |