Rui Xie 0003

dblp:86/2228-3 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-1756-7746ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 9 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation
abstract
Multilingual Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to perform knowledge-intensive tasks in multilingual settings by leveraging retrieved documents as external evidence. However, when the retrieved evidence differs in language from the user query and in-context exemplars, the model often exhibits language drift by generating responses in an unintended language. This phenomenon is especially pronounced during reasoning-intensive decoding, such as Chain-of-Thought (CoT) generation, where intermediate steps introduce further language instability. In this paper, we systematically study output language drift in multilingual RAG across multiple datasets, languages, and LLM backbones. Our controlled experiments reveal that the drift results not from comprehension failure but from decoder-level collapse, where dominant token distributions and high-frequency English patterns dominate the intended generation language. We further observe that English serves as a semantic attractor under cross-lingual conditions, emerging as both the strongest interference source and the most frequent fallback language. To mitigate this, we propose Soft Constrained Decoding (SCD), a lightweight, training-free decoding strategy that gently steers generation toward the target language by penalizing non-target-language tokens. SCD is model-agnostic and can be applied to any generation algorithm without modifying the architecture or requiring additional data. Experiments across three multilingual datasets and multiple typologically diverse languages show that SCD consistently improves language alignment and task performance, providing an effective and generalizable solution in multilingual RAG.
Bo Li 0099, Zhenghua Xu 0001, Rui Xie 0003
AAAI3
2026 L2-LoRA: Improving Low-Rank Adaptation with Layer-Specific Regularization
abstract
Fine-tuning large language models (LLMs) in a parameter-efficient manner while preserving their pre-trained world knowledge remains a significant challenge. While Low-Rank Adaptation (LoRA) and its variants effectively mitigate catastrophic forgetting, they do not fully eliminate the loss of critical pre-trained knowledge. In this work, we first analyze the layer-wise distribution of domain-specific knowledge within LLMs through knowledge localization, and empirically identify a clear layer-specific pattern: pre-trained world knowledge predominantly resides in lower layers, whereas knowledge relevant to downstream tasks is more concentrated in higher layers. Motivated by this observation, we propose L2-LoRA, a simple yet effective variant of LoRA that applies layer-specific L2 regularization to the LoRA weights during fine-tuning. Specifically, L2-LoRA imposes stronger regularization on lower layers to preserve pre-trained world knowledge, while allowing greater adaptation in higher layers to better align with downstream tasks. Experiments across multiple benchmarks show that L2-LoRA not only consistently outperforms vanilla LoRA in downstream performance, but also effectively mitigates catastrophic forgetting by retaining more pre-trained knowledge.
Rui Xie 0003, Shikun Zhang
AAAI2
2026 An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
Chengli Xing, Zhengran Zeng, Gexiang Fang, Rui Xie 0003, Wei Ye 0004, Shikun Zhang
ICPC4
2025 All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image Denoising
abstract
Image denoising poses a significant challenge in image processing, aiming to remove noise and artifacts from input images. However, current denoising algorithms implemented on electronic chips frequently encounter latency issues and demand substantial computational resources. In this paper, we introduce an all-optical Nonlinear Diffractive Denoising Deep Network (N3DNet) for image denoising at the speed of light. Initially, we incorporate an image encoding and pre-denoising module into the Diffractive Deep Neural Network and integrate a nonlinear activation function, termed the phase exponential linear function, after each diffractive layer, thereby boosting the network’s nonlinear modeling and denoising capabilities. Subsequently, we devise a new reinforcement learning algorithm called regularization-assisted deep Q-network to optimize N3DNet. Finally, leveraging 3D printing techniques, we fabricate N3DNet using the trained parameters and construct a physical experimental system for real-world applications. A new benchmark dataset, termed MIDD, is constructed for mode image denoising, comprising 120K pairs of noisy/noise-free images captured from real fiber communication systems across various transmission lengths. Through extensive simulation and real experiments, we validate that N3DNet outperforms both traditional and deep learning-based denoising approaches across various datasets. Remarkably, its processing speed is nearly 3,800 times faster than electronic chip-based methods.
Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Rui Xie 0003, Guanju Peng, Shikun Zhang
CVPR4
2025 Mitigating spurious correlations with causal logit perturbation
Xiaoling Zhou, Wei Ye 0004, Rui Xie 0003, Shikun Zhang
Inf. Sci.3
2024 Enhancing In-Context Learning via Implicit Demonstration Augmentation
abstract
Xiaoling Zhou, Wei Ye, Yidong Wang, Chaoya Jiang, Zhemg Lee, Rui Xie, Shikun Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaoling Zhou, Wei Ye 0004, Yidong Wang 0003, Chaoya Jiang, Zhemg Lee, Rui Xie 0003, Shikun Zhang
ACL (1)6
2024 PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
abstract
Instruction tuning large language models (LLMs) remains a challenging task, owing to the complexity of hyperparameter selection and the difficulty involved in evaluating the tuned models. To determine the optimal hyperparameters, an automatic, robust, and reliable evaluation benchmark is essential. However, establishing such a benchmark is not a trivial task due to the challenges associated with evaluation accuracy and privacy protection. In response to these challenges, we introduce a judge large language model, named PandaLM, which is trained to distinguish the superior model given several LLMs. PandaLM's focus extends beyond just the objective correctness of responses, which is the main focus of traditional evaluation datasets. It addresses vital subjective factors such as relative conciseness, clarity, adherence to instructions, comprehensiveness, and formality. To ensure the reliability of PandaLM, we collect a diverse human-annotated test dataset, where all contexts are generated by humans and labels are aligned with human preferences. Our findings reveal that PandaLM-7B offers a performance comparable to both GPT-3.5 and GPT-4. Impressively, PandaLM-70B surpasses their performance. PandaLM enables the evaluation of LLM to be fairer but with less cost, evidenced by significant improvements achieved by models tuned through PandaLM compared to their counterparts trained with default Alpaca's hyperparameters. In addition, PandaLM does not depend on API-based evaluations, thus avoiding potential data leakage.
Yidong Wang 0003, Zhuohao Yu 0001, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen 0102, Chaoya Jiang, Rui Xie 0003, Jindong Wang 0001, Xing Xie 0001, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004
ICLR9
2024 Improving Long-Tail Vulnerability Detection Through Data Augmentation Based on Large Language Models
abstract
The ability of automatic vulnerability detection models largely depends on the dataset used for training. However, annotating these datasets is costly and time-consuming, leading to a scarcity of labeled samples, particularly for diverse CWE types. This scarcity results in a pronounced long-tail issue, where less common types are underrepresented, thus diminishing the model's effectiveness in detecting them. In this paper, we address this challenge by employing large language models' (LLMs) generative and reasoning capabilities to create the necessary training samples for detecting less common vulnerability types. Specifically, we use GPT-4 to generate targeted samples of these types. After a well-defined self-filtering process, these samples are incorporated into the training of detection models. Extensive experiments show that our approach significantly enhances vulnerability detection capabilities, especially for long-tail vulnerabilities, across a variety of detection models, including both traditional deep learning and modern LLM-based detection models. Further, our comparative tests on unseen projects indicate that models trained with our generated data can identify more real-world vulnerabilities than traditional methods, proving the practicality and generalizability of our approach in real settings. The code and the generated dataset are publicly available at https://github.com/LuckyDengXiao/LERT.
Fuyao Duan, Rui Xie 0003, Wei Ye 0004, Shikun Zhang
ICSME3
2024 Boosting Model Resilience via Implicit Adversarial Data Augmentation
Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Rui Xie 0003, Shikun Zhang
IJCAI4
2024 Leveraging In-and-Cross Project Pseudo-Summaries for Project-Specific Code Summarization
abstract
Code summarization is pivotal in software development, aiding developers in grasping the semantics of source code. However, existing research predominantly focuses on the general code summarization capabilities of models, neglecting project-specific summary characteristics. However, given the scarcity of project-internal code summary corpora, enhancing the model’s performance for a specific project presents a significant challenge. To tackle this issue, we introduces the use of In-and-Cross project pseudo-summaries to improve Project-Specific Code Summarization. Specifically, we employ models trained on other projects to generate cross-project pseudo-summaries and learn the distinctions from target-project through contrastive learning. Simultaneously, we utilize in-project pseudo-summaries generated by the current model, harnessing these data through semi-supervised learning to enhance performance. The experiment results show that the proposed method can effectively improve the performance of the summarization task in practical scenarios, and can also enhance the coordination of the model.
Tianxiang Hu, Ninglin Liao, Rui Xie 0003, Dongdong Du, Shujun Lin
IJCNN4
2024 CoderUJB: An Executable and Unified Java Benchmark for Practical Programming Scenarios
abstract
In the evolving landscape of large language models (LLMs) tailored for software engineering, the need for benchmarks that accurately reflect real-world development scenarios is paramount. Current benchmarks are either too simplistic or fail to capture the multi-tasking nature of software development. To address this, we introduce CoderUJB, a new benchmark designed to evaluate LLMs across diverse Java programming tasks that are executable and reflective of actual development scenarios, acknowledging Java's prevalence in real-world software production. CoderUJB comprises 2,239 programming questions derived from 17 real open-source Java projects and spans five practical programming tasks. Our empirical study on this benchmark investigates the coding abilities of various open-source and closed-source LLMs, examining the effects of continued pre-training in specific programming languages code and instruction fine-tuning on their performance. The findings indicate that while LLMs exhibit strong potential, challenges remain, particularly in non-functional code generation (e.g., test generation and defect detection). Importantly, our results advise caution in the specific programming languages continued pre-training and instruction fine-tuning, as these techniques could hinder model performance on certain tasks, suggesting the need for more nuanced strategies. CoderUJB thus marks a significant step towards more realistic evaluations of programming capabilities in LLMs, and our study provides valuable insights for the future development of these models in software engineering.
Zhengran Zeng, Yidong Wang 0003, Rui Xie 0003, Wei Ye 0004, Shikun Zhang
ISSTA3
2024 Exploring Vision-Language Models for Imbalanced Learning
Yidong Wang 0003, Zhuohao Yu 0001, Jindong Wang 0001, Qiang Heng, Hao Chen 0102, Wei Ye 0004, Rui Xie 0003, Xing Xie 0001, Shikun Zhang
Int. J. Comput. Vis.7
2022 3Rs: Data Augmentation Techniques Using Document Contexts For Low-Resource Chinese Named Entity Recognition
abstract
With recent advances of neural networks and pre-training techniques, Chinese Named Entity Recognition (NER) has achieved great progress in recent years. However, NER systems still have the problem of generalization ability issues due to lack of annotated data, and current NER models mostly consider input sentences individually, which prevent models from further exploiting cross-sentence document context in training. With regard of these problems, this paper present new insights into Chinese NER and propose 3Rs: three data augmentation methods incorporating document-level information for NER through random concatenating, random swapping and random erasing, which are inspired by some multi-sample data augmentation techniques in computer vision fields, aiming to reorganize the composition of training sentences, and generate more training examples with less human efforts. We conduct extensive experiments on two Chinese datasets, and introduce a two-level attacking method to audit robustness performance. Our experiment results show that even the best model can obtain a better accuracy and robustness, especially for smaller training sets, therefore alleviating performance bottlenecks on low-resource conditions.
Zheyu Ying, Rui Xie 0003, Guochang Wen, Xueyang Liu, Shikun Zhang
IJCNN3
2022 Low-Resources Project-Specific Code Summarization
abstract
Code summarization generates brief natural language descriptions of source code pieces, which can assist developers in understanding code and reduce documentation workload. Recent neural models on code summarization are trained and evaluated on large-scale multi-project datasets consisting of independent code-summary pairs. Despite the technical advances, their effectiveness on a specific project is rarely explored. In practical scenarios, however, developers are more concerned with generating high-quality summaries for their working projects. And these projects may not maintain sufficient documentation, hence having few historical code-summary pairs. To this end, we investigate low-resource project-specific code summarization, a novel task more consistent with the developers’ requirements. To better characterize project-specific knowledge with limited training samples, we propose a meta transfer learning method by incorporating a lightweight fine-tuning mechanism into a meta-learning framework. Experimental results on nine real-world projects verify the superiority of our method over alternative ones and reveal how the project-specific knowledge is learned.
Rui Xie 0003, Tianxiang Hu, Wei Ye 0004, Shikun Zhang
ASE1
2021 Exploiting Method Names to Improve Code Summarization: A Deliberation Multi-Task Learning Approach
abstract
Code summaries are brief natural language descriptions of source code pieces. The main purpose of code summarization is to assist developers in understanding code and to reduce documentation workload. In this paper, we design a novel multi-task learning (MTL) approach for code summarization through mining the relationship between method code summaries and method names. More specifically, since a method's name can be considered as a shorter version of its code summary, we first introduce the tasks of generation and informativeness prediction of method names as two auxiliary training objectives for code summarization. A novel two-pass deliberation mechanism is then incorporated into our MTL architecture to generate more consistent intermediate states fed into a summary decoder, especially when informative method names do not exist. To evaluate our deliberation MTL approach, we carried out a large-scale experiment on two existing datasets for Java and Python. The experiment results show that our technique can be easily applied to many state-of-the-art neural models for code summarization and improve their performance. Meanwhile, our approach shows significant superiority when generating summaries for methods with non-informative names.
Rui Xie 0003, Wei Ye 0004, Jinan Sun, Shikun Zhang
ICPC1
2020 Graph Enhanced Dual Attention Network for Document-Level Relation Extraction
abstract
Document-level relation extraction requires inter-sentence reasoning capabilities to capture local and global contextual information for multiple relational facts.To improve inter-sentence reasoning, we propose to characterize the complex interaction between sentences and potential relation instances via a Graph Enhanced Dual Attention network (GEDA).In GEDA, sentence representation generated by the sentence-to-relation (S2R) attention is refined and synthesized by a Heterogeneous Graph Convolutional Network before being fed into the relation-to-sentence (R2S) attention .We further design a simple yet effective regularizer based on the natural duality of the S2R and R2S attention, whose weights are also supervised by the supporting evidence of relation instances during training.An extensive set of experiments on an existing large-scale dataset show that our model achieves competitive performance, especially for the inter-sentence relation extraction, while the neural predictions can also be interpretable and easily observed.
Bo Li 0099, Wei Ye 0004, Zhonghao Sheng, Rui Xie 0003, Xiangyu Xi, Shikun Zhang
COLING4
2020 Exploiting Code Knowledge Graph for Bug Localization via Bi-directional Attention
abstract
Bug localization automatic localize relevant source files given a natural language description of bug within a software project. For a large project containing hundreds and thousands of source files, developers need cost lots of time to understand bug reports generated by quality assurance and localize these buggy source files. Traditional methods are heavily depending on the information retrieval technologies which rank the similarity between source files and bug reports in lexical level. Recently, deep learning based models are used to extract semantic information of code with significant improvements for bug localization. However, programming language is a highly structural and logical language, which contains various relations within and cross source files. Thus, we propose KGBugLocator to utilize knowledge graph embeddings to extract these interrelations of code, and a keywords supervised bi-directional attention mechanism regularize model with interactive information between source files and bug reports. With extensive experiments on four different projects, we prove our model can reach the new the-state-of-art(SOTA) for bug localization.
Rui Xie 0003, Wei Ye 0004, Shikun Zhang
ICPC2
2020 Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual Learning
abstract
Code summarization generates brief natural language description given a source code snippet, while code retrieval fetches relevant source code given a natural language query. Since both tasks aim to model the association between natural language and programming language, recent studies have combined these two tasks to improve their performance. However, researchers have yet been able to effectively leverage the intrinsic connection between the two tasks as they train these tasks in a separate or pipeline manner, which means their performance can not be well balanced. In this paper, we propose a novel end-to-end model for the two tasks by introducing an additional code generation task. More specifically, we explicitly exploit the probabilistic correlation between code summarization and code generation with dual learning, and utilize the two encoders for code summarization and code generation to train the code retrieval task via multi-task learning. We have carried out extensive experiments on an existing dataset of SQL and Python, and results show that our model can significantly improve the results of the code retrieval task over the-state-of-art models, as well as achieve competitive performance in terms of BLEU score for the code summarization task.
Wei Ye 0004, Rui Xie 0003, Tianxiang Hu, Xiaoyin Wang, Shikun Zhang
WWW2
2019 Exploiting Entity BIO Tag Embeddings and Multi-task Learning for Relation Extraction with Imbalanced Data
abstract
In practical scenario, relation extraction needs to first identify entity pairs that have relation and then assign a correct relation class.However, the number of non-relation entity pairs in context (negative instances) usually far exceeds the others (positive instances), which negatively affects a model's performance.To mitigate this problem, we propose a multitask architecture which jointly trains a model to perform relation identification with crossentropy loss and relation classification with ranking loss.Meanwhile, we observe that a sentence may have multiple entities and relation mentions, and the patterns in which the entities appear in a sentence may contain useful semantic information that can be utilized to distinguish between positive and negative instances.Thus we further incorporate the embeddings of character-wise/word-wise BIO tag from the named entity recognition task into character/word embeddings to enrich the input representation.Experiment results show that our proposed approach can significantly improve the performance of a baseline model with more than 10% absolute increase in F1-score, and outperform the state-of-theart models on ACE 2005 Chinese and English corpus.Moreover, BIO tag embeddings are particularly effective and can be used to improve other models as well.* indicates equal contribution.
Wei Ye 0004, Bo Li 0099, Rui Xie 0003, Zhonghao Sheng, Shikun Zhang
ACL (1)3
2019 A Hybrid Character Representation for Chinese Event Detection
abstract
For the Chinese language, event triggers in a sentence may appear inside or across words after word segmentation. Thus recent works on Chinese event detection often formulate the task as a character-wise sequence labeling problem instead of a word-wise one. Due to a limited amount of corpus, however, it is more difficult in practice to train character-wise models to capture the inner structure of event triggers and the semantics of sentence-level context compared with word-wise ones. In this paper, we propose to improve character-wise models by incorporating word information and language model representation into Chinese character representation. More specifically, the former consists of the position of the character inside a word and the word's embedding, which can aid structural pattern learning; the latter is obtained by BERT, which contains long-distance semantic information. We construct a sequence tagging model equipped with the hybrid representation and evaluate our model on ACE 2005 Chinese corpus. Experiment results show that both word information and language model representation are effective enhancements, and our model gains an increase of 4.5 (6.5%) and 6.1 (9.4%) in F1-score in event trigger identification task and classification task respectively over the state-of-the-art method.
Xiangyu Xi, Tong Zhang 0001, Wei Ye 0004, Rui Xie 0003, Shikun Zhang
IJCNN5
2019 DeepLink: A Code Knowledge Graph Based Deep Learning Approach for Issue-Commit Link Recovery
abstract
Links between issue reports and corresponding code commits to fix them can greatly reduce the maintenance costs of a software project. More often than not, however, these links are missing and thus cannot be fully utilized by developers. Current practices in issue-commit link recovery extract text features and code features in terms of textual similarity from issue reports and commit logs to train their models. These approaches are limited since semantic information could be lost. Furthermore, few of them consider the effect of source code files related to a commit on issue-commit link recovery, let alone the semantics of code context. To tackle these problems, we propose to construct code knowledge graph of a code repository and generate embeddings of source code files to capture the semantics of code context. We also use embeddings to capture the semantics of issue- or commit-related text. Then we use these embeddings to calculate semantic similarity and code similarity using a deep learning approach before training a SVM binary classification model with additional features. Evaluations on real-world projects show that our approach DeepLink can outperform the state-of-the-art method.
Rui Xie 0003, Wei Ye 0004, Tianxiang Hu, Dongdong Du, Shikun Zhang
SANER1