VLDB 2026 Research / reviewers in the wild / expert
Xilin Liu 0001
dblp:05/8005-1
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0001-4870-1012ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) has proven effective in integrating external knowledge into large language models (LLMs) for solving question-answer (QA) tasks. The state-of-the-art RAG approaches often use the graph data as the external data since they capture the rich semantic information and link relationships between entities. However, existing graph-based RAG approaches cannot accurately identify the relevant information from the graph and also consume large numbers of tokens in the online retrieval process. To address these issues, we introduce a novel graph-based RAG approach, called Attributed Community-based Hierarchical RAG (ArchRAG), by augmenting the question using attributed communities, and also introducing a novel LLM-based hierarchical clustering method. To retrieve the most relevant information from the graph for the question, we build a novel hierarchical index structure for the attributed communities and develop an effective online retrieval method. Experimental results demonstrate that ArchRAG outperforms existing methods in both accuracy and token cost. Yixiang Fang, Yingli Zhou, Xilin Liu 0001, Yuchi Ma |
AAAI | 4 |
| 2026 | EffiReasonTrans: RL-Optimized Reasoning for Code TranslationabstractCode translation is a crucial task in software development and maintenance. While recent advancements in Large Language Models (LLMs) have improved automated code translation accuracy, these gains often come at the cost of increased inference latency–hindering real-world development workflows that involve human-in-the-loop inspection. To address this tradeoff, we propose EffiReasonTrans, a training framework designed to improve translation accuracy while balancing inference latency. We first construct a high-quality reasoning-augmented dataset by prompting a stronger language model DeepSeek-R1 to generate intermediate reasoning and target translations. Each (source code, reasoning, target code) triplet undergoes automated syntax and functionality checks to ensure reliability. Based on this dataset, we employ a two-stage training strategy: supervised fine-tuning on reasoning-augmented samples, followed by reinforcement learning to further enhance accuracy, which also helps balance inference latency. We evaluate EffiReasonTrans on six translation pairs. Experimental results show that EffiReason-Trans consistently improves translation accuracy (up to +49.2% CA and +27.8% CodeBLEU compared to the base model), while reducing the number of generated tokens (up to -19.3%) and lowering inference latency in most cases (up to -29.0%). Ablation studies further confirm the complementary benefits of the two-stage training framework. Additionally, EffiReasonTrans shows improvements of translation accuracy when integrated into agent-based frameworks. Our code and data are available athttps://github.com/DeepSoftwareAnalytics/EffiReasonTrans. Yanlin Wang 0001, Rongyi Ou, Yanli Wang 0001, Mingwei Liu 0002, Jiachi Chen, Ensheng Shi, Xilin Liu 0001, Yuchi Ma, Zibin Zheng |
IEEE Trans. Software Eng. | 7 |
| 2026 | RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code TranslationabstractRepository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source repository. Many benchmarks have been proposed to evaluate the performance of such code translators. However, previous benchmarks mostly provide fine-grained samples, focusing at either code snippet, function, or file-level code translation. Such benchmarks do not accurately reflect real-world demands, where entire repositories often need to be translated, involving longer code length and more complex functionalities. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code translation benchmark featuring 1,897 real-world repository samples across 13 language pairs with automatically executable test suites. Besides, we introduce RepoTransAgent, a general agent framework to perform repository-level code translation. We evaluate both our benchmark’s challenges and agent’s effectiveness using several methods and backbone LLMs, revealing that repository-level translation remains challenging, where the best-performing method achieves only a 32.8% success rate. Furthermore, our analysis reveals that translation difficulty varies significantly by language pair direction, with dynamic-to-static language translation being much more challenging than the reverse direction (achieving below 10% vs. static-to-dynamic at 45-63%). Finally, we conduct a detailed error analysis and highlight current LLMs’ deficiencies in repository-level code translation, which could provide a reference for further improvements. We provide the code and data athttps://github.com/DeepSoftwareAnalytics/RepoTransBench. Yanli Wang 0001, Yanlin Wang 0001, Suiquan Wang, Daya Guo, Jiachi Chen, John C. Grundy, Xilin Liu 0001, Yuchi Ma, Mingzhi Mao, Hongyu Zhang 0002, Zibin Zheng |
IEEE Trans. Software Eng. | 7 |
| 2025 | DrainCode: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context PoisoningabstractLarge language models (LLMs) have demonstrated impressive capabilities in code generation, by leveraging retrieval-augmented generation (RAG) methods. However, the computational costs associated with LLM inference, particularly in terms of latency and energy consumption, have received limited attention in the security context. This paper introduces DrainCode, the first adversarial attack targeting the computational efficiency of RAG-based code generation systems. By strategically poisoning retrieval contexts through mutation-based approach, DrainCode forces LLMs to produce significantly longer outputs, thereby increasing GPU latency and energy consumption. We evaluate the effectiveness of DrainCode across multiple models. Our experiments show that DrainCode achieves up to a 85% increase in latency, a 49% increase in energy consumption, and more than a 3× increase in output length compared to the baseline. Furthermore, we demonstrate the generalizability of the attack across different prompting strategies and its effectiveness compared to different defenses. The results highlight DrainCode as a potential method for increasing the computational overhead of LLMs, making it useful for evaluating LLM security in resource-constrained environments. We provide code and data at https://github.com/DeepSoftwareAnalytics/DrainCode. Yanli Wang 0001, Jiadong Wu, Tianyue Jiang, Mingwei Liu 0002, Jiachi Chen, Chong Wang 0013, Ensheng Shi, Xilin Liu 0001, Yuchi Ma, Zibin Zheng |
ASE | 8 |
| 2025 | In-depth Analysis of Graph-based RAG in a Unified Framework
Yingli Zhou, Yaodong Su, Youran Sun, Taotao Wang, Runyuan He, Sicong Liang, Xilin Liu 0001, Yuchi Ma, Yixiang Fang |
Proc. VLDB Endow. | 9 |
| 2024 | On Efficient Large Sparse Matrix Chain MultiplicationabstractSparse matrices are often used to model the interactions among different objects and they are prevalent in many areas including e-commerce, social network, and biology. As one of the fundamental matrix operations, the sparse matrix chain multiplication (SMCM) aims to efficiently multiply a chain of sparse matrices, which has found various real-world applications in areas like network analysis, data mining, and machine learning. The efficiency of SMCM largely hinges on the order of multiplying the matrices, which further relies on the accurate estimation of the sparsity values of intermediate matrices. Existing matrix sparsity estimators often struggle with large sparse matrices, because they suffer from the accuracy issue in both theory and practice. To enable efficient SMCM, in this paper we introduce a novel row-wise sparsity estimator (RS-estimator), a straightforward yet effective estimator that leverages matrix structural properties to achieve efficient, accurate, and theoretically guaranteed sparsity estimation. Based on the RS-estimator, we propose a novel ordering algorithm for determining a good order of efficient SMCM. We further develop an efficient parallel SMCM algorithm by effectively utilizing multiple CPU threads. We have conducted experiments by multiplying various chains of large sparse matrices extracted from five real-world large graph datasets, and the results demonstrate the effectiveness and efficiency of our proposed methods. In particular, our SMCM algorithm is up to three orders of magnitude faster than the state-of-the-art algorithms. Chunxu Lin, Wensheng Luo 0002, Yixiang Fang, Chenhao Ma 0001, Xilin Liu 0001, Yuchi Ma |
Proc. ACM Manag. Data | 5 |