Rui Li 0093

dblp:96/4282-93 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0009-0005-3657-1133ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 From Diagnosis to Generalization: A Cognitive Approach to Data Selection for Educational LLMs
abstract
Specializing Large Language Models for educational domains is a key frontier in creating personalized learning tools. The central challenge is not data scarcity but its abundance: efficiently selecting a curated data subset from vast corpora to enhance specialized skills and foster generalization, without degrading existing abilities. Existing data selection paradigms, relying on superficial semantic similarity or model training dynamics, often lack a principled framework to identify data that promotes true cognitive growth. Our work proposes a paradigm shift from leveraging indirect proxies of learning value, such as semantic similarity and training dynamics, towards a framework that performs a direct, cognitive-level modeling of the learner's state. We introduce CASS, a novel framework that implements this cognitive approach through a clear pipeline, moving from an initial Diagnosis to the ultimate goal of expanding the model's cognitive frontier. First, CASS diagnoses the LLM's cognitive frontier using Multidimensional Item Response Theory. Leveraging this diagnosis, it then employs Fisher Information to select a data subset situated at LLM's cognitive frontier that offers maximum informational gain. Finally, the model is fine-tuned on this curated data using a structured, easy-to-hard curriculum to ensure effective learning. Experiments on our new multi-subject dataset show that models trained with CASS not only achieve superior accuracy in the target domain but also exhibit enhanced generalization. CASS provides a more efficient, effective, and theoretically-grounded paradigm for building expert educational LLMs.
Yuxiang Guo 0002, Yan Zhuang 0001, Qi Liu 0003, Zhenya Huang, Xianquan Wang, Liyang He, Jiatong Li 0002, Rui Li 0093, Shijin Wang 0001
AAAI8
2026 LongTutor: Benchmarking Large Language Models for Long-term Personalized Tutoring
abstract
Ning Li, Zheng Zhang, Zhenya Huang, Rui Li, Yi Zhan, Yinbo Luo, Qi Liu, Enhong Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ning Li 0055, Zheng Zhang 0048, Zhenya Huang, Rui Li 0093, Yinbo Luo, Qi Liu 0003, Enhong Chen
ACL (1)4
2026 Controllable Contamination Detection for Reliable LLM Evaluation with Statistical Guarantees
abstract
Zheng Zhang, Qi Liu, Siyuan Liang, Ning Li, Zirui Hu, Weibo Gao, Rui Li, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zheng Zhang 0048, Qi Liu 0003, Siyuan Liang 0004, Ning Li 0055, Zirui Hu, Weibo Gao, Rui Li 0093, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao
ACL (1)7
2026 PoTable: Toward Systematic Thinking via Plan-Then-Execute Stage Reasoning on Tables
abstract
In recent years, table reasoning has garnered substantial research interest, particularly regarding its integration with Large Language Models (LLMs), which have revolutionized natural language applications. Existing LLM-based studies typically achieve step-by-step thinking for table reasoning guided by task semantics. While these approaches emphasize autonomous exploration and enhance fine-grained table understanding, they often overlook systematic thinking in the reasoning process. This oversight can lead to omitted steps, disorganized logic and misleading results, especially in complex scenarios. In this paper, we proposePoTable, a novel stage-oriented plan-then-execute approach that incorporates systematic thinking into table reasoning. Specifically,PoTableinvolves several distinct analytical stages with clear objectives to provide adequate guidance. To accomplish stage-specific goals,PoTableemploys a plan-then-execute mechanism: it first plans the operation chain based on the stage objective, and then executes operations sequentially through code generation, real-time running and feedback processing. Consequently,PoTableproduces reliable table reasoning results with highly accurate, step-wise commented and completely executable programs. It mirrors the workflow of a professional data analyst, offering advantages in both accuracy and explainability. Finally, we conduct extensive experiments on four datasets from the WikiTQ and TabFact benchmarks, where the results demonstrate the effectiveness, efficiency and explainability ofPoTable. Our code is available at:https://github.com/Double680/PoTable.
Qingyang Mao, Qi Liu 0003, Zhi Li 0057, Mingyue Cheng 0004, Zheng Zhang 0048, Rui Li 0093
IEEE Trans. Knowl. Data Eng.6
2026 The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant document from external knowledge sources. By referencing this external knowledge, RAG effectively reduces the generation of factually incorrect content and addresses hallucination issues within LLMs. Recently, there has been growing attention to improving the performance and efficiency of RAG systems from various perspectives. While these advancements have yielded significant results, the application of RAG in domains with considerable societal implications raises a critical question about fairness: What impact does the introduction of the RAG paradigm have on the fairness of LLMs? To address this question, we conduct extensive experiments by varying the LLMs, retrievers, and retrieval sources. Our experimental analysis reveals that the scale of the LLMs plays a significant role in influencing fairness outcomes within the RAG framework. When the model scale is smaller than 8B, the integration of retrieval mechanisms often exacerbates unfairness in small-scale LLMs (e.g., LLaMA3.2-1B, Mistral-7B, and LLaMA3-8B). To mitigate the fairness issues introduced by RAG for small-scale LLMs, we propose two approaches, FairFT and FairFilter. Specifically, in FairFT, we align the retriever with the LLM in terms of fairness, enabling it to retrieve documents that facilitate fairer model outputs. In FairFilter, we propose a fairness filtering mechanism to filter out biased content after retrieval. Finally, we validate our proposed approaches on real-world datasets, demonstrating their effectiveness in improving fairness while maintaining performance.
Zheng Zhang 0048, Ning Li 0055, Qi Liu 0003, Rui Li 0093, Weibo Gao, Qingyang Mao, Zhenya Huang, Baosheng Yu, Dacheng Tao
IEEE Trans. Knowl. Data Eng.4
2025 VERSE: Verification-based Self-Play for Code Instructions
abstract
Instruction-tuned Code Large Language Models (Code LLMs) have excelled in diverse code-related tasks, such as program synthesis, automatic program repair, and code explanation. To collect training datasets for instruction-tuning, a popular method involves having models autonomously generate instructions and corresponding responses. However, the direct generation of responses does not ensure functional correctness, a crucial requirement for generating responses to code instructions. To overcome this, we present Verification-Based Self-Play (VERSE), aiming to enhance model proficiency in generating correct responses. VERSE establishes a robust verification framework that covers various code instructions. Employing VERSE, Code LLMs engage in self-play to generate instructions and corresponding verifications. They evaluate execution results and self-consistency as verification outcomes, using them as scores to rank generated data for self-training. Experiments show that VERSE improves multiple base Code LLMs (average 7.6%) across various languages and tasks on many benchmarks, affirming its effectiveness.
Hao Jiang 0023, Qi Liu 0003, Rui Li 0093, Yuze Zhao, Shengyu Ye, Junyu Lu 0003, Yu Su 0002
AAAI3
2025 Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document Relationship
abstract
Dense retrieval has emerged as the leading approach in information retrieval, aiming to find semantically relevant documents based on natural language queries. Given that a single document can be retrieved by multiple distinct queries, existing methods aim to represent a document with multiple vectors. Each vector is aligned with a different query to model the many-to-one relationship between queries and documents. However, these multiple vector-based approaches encounter challenges such as Increased Storage, Vector Collapse, and Search Efficiency. To address these issues, we introduce the Distribution-Driven Dense Retrieval framework (DDR). Specifically, we use vectors to represent queries and distributions to represent documents. This approach not only captures the relationships between multiple queries corresponding to the same document but also avoids the need to use multiple vectors to represent the document. Furthermore, to ensure search efficiency for DDR, we propose a dot product-based computation method to calculate the similarity between documents represented by distributions and queries represented by vectors. This allows for seamless integration with existing approximate nearest neighbor (ANN) search algorithms for efficient search. Finally, we conduct extensive experiments on real-world datasets, which demonstrate that our method significantly outperforms traditional dense retrieval methods.
Junfeng Kang, Rui Li 0093, Qi Liu 0003, Zhenya Huang, Zheng Zhang 0048, Yanjiang Chen, Linbo Zhu, Yu Su 0002
AAAI2
2025 PQR: Improving Dense Retrieval via Potential Query Modeling
abstract
Dense retrieval has now become the mainstream paradigm in information retrieval. The core idea of dense retrieval is to align document embeddings with their corresponding query embeddings by maximizing their dot product. The current training data is quite sparse, with each document typically associated with only one or a few labeled queries. However, a single document can be retrieved by multiple different queries. Aligning a document with just one or a limited number of labeled queries results in a loss of its semantic information. In this paper, we propose a training-free Potential Query Retrieval (PQR) framework to address this issue. Specifically, we use a Gaussian mixture distribution to model all potential queries for a document, aiming to capture its comprehensive semantic information. To obtain this distribution, we introduce three sampling strategies to sample a large number of potential queries for each document and encode them into a semantic space. Using these sampled queries, we employ the Expectation-Maximization algorithm to estimate parameters of the distribution. Finally, we also propose a method to calculate similarity scores between user queries and documents under the PQR framework. Extensive experiments demonstrate the effectiveness of the proposed method.
Junfeng Kang, Rui Li 0093, Qi Liu 0003, Yanjiang Chen, Zheng Zhang 0048, Junzhe Jiang 0001, Yu Su 0002
ACL (1)2
2025 UniRAG: Unified Query Understanding Method for Retrieval Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) technology effectively addresses the issues of knowledge update lag and hallucinations in large language models (LLMs) by integrating internal and external knowledge. Existing query augmentation methods improve RAG’s performance in handling complex queries but face two key challenges: (1) the separation of query augmentation and encoding tasks, which hinders information sharing and introduces cumulative errors, and (2) the difficulty of selecting the optimal augmentation strategy for different scenarios. In this work, we propose UniRAG, a unified framework for query understanding in RAG. UniRAG employs a decoder-only LLM to jointly perform query augmentation and encoding, eliminating task separation. To facilitate adaptive query augmentation, we categorize existing techniques into query paraphrasing, query expansion, and query abstraction. Our model learns to select the optimal augmentation strategy based on user queries, leveraging retrieval and generation outputs as feedback. Experimental results show that UniRAG significantly outperforms traditional query augmentation methods in five knowledge-intensive benchmark tasks in both closed and open domain question answering.
Rui Li 0093, Liyang He, Qi Liu 0003, Zheng Zhang 0048, Yuyang Ye 0002, Linbo Zhu, Yu Su 0002
ACL (1)1
2025 CursorCore: Assist Programming through Aligning Anything
abstract
Large language models have been successfully applied to programming assistance tasks, such as code completion, code insertion, and instructional code editing. However, these applications remain insufficiently automated and struggle to effectively integrate various types of information during the programming process, including coding history, code context, and user instructions. In this work, we propose a new framework that comprehensively integrates these information sources, collect data to train our models and evaluate their performance. Firstly, to thoroughly evaluate how well models align with different types of information and the quality of their outputs, we introduce a new benchmark, APEval (Assist Programming Eval), to comprehensively assess the performance of models in programming assistance tasks. Then, for data collection, we develop a data generation pipeline, Programming-Instruct, which synthesizes training data from diverse sources, such as GitHub and online judge platforms. This pipeline can automatically generate various types of messages throughout the programming process. Finally, using this pipeline, we generate 219K samples, fine-tune multiple models, and develop the CursorCore series. We show that CursorCore outperforms other models of comparable size. This framework unifies applications such as inline chat and automated editing, contributes to the advancement of coding assistants.
Hao Jiang 0023, Qi Liu 0003, Rui Li 0093, Shengyu Ye, Shijin Wang 0001
ICML3
2025 Denoising Programming Knowledge Tracing with a Code Graph-based Tuning Adaptor
abstract
Programming Knowledge Tracking (PKT) aims to dynamically diagnose learners' mastery levels of programming knowledge based on their coding activities, facilitating more effective and personalized programming education. However, current PKT studies primarily focus on the implicit relationship between code content and knowledge assessment, often overlooking two types of noise signals in long-term programming activities: unwanted signals from unrelated submissions and weak signals from minor modifications. This practical challenge significantly limits model performance and application. To address this issue, we propose Coda, a Code graph-based tuning adaptor designed to enhance existing PKT models by identifying and mitigating the impact of noise. Specifically, Coda first transforms the loose code sequences submitted by each learner into a compact code graph. By leveraging this code graph, unwanted signals can be identified from a semantic similarity perspective. We then apply a cluster-aware GCN to the code graph, which improves the discrimination of weak signals and enables their clustering for identification. Finally, a lightweight yet effective adaptor is incorporated into the PKT task through optimization with two noise feature-based constraints and a navigational regularization term, to correct knowledge states affected by noise. It is worth mentioning that the Coda framework is model-agnostic and can be adapted to most existing PKT solutions. Extensive experimental results on four real-world datasets demonstrate that Coda effectively performs the PKT task in the presence of noisy programming records, outperforming typical baselines.
Weibo Gao, Qi Liu 0003, Rui Li 0093, Yuze Zhao, Hao Wang 0076, Linan Yue, Fangzhou Yao, Zheng Zhang 0048
KDD (1)3
2025 MGS3: A Multi-Granularity Self-Supervised Code Search Framework
abstract
In the pursuit of enhancing software reusability and developer productivity, code search has emerged as a key area, aimed at retrieving code snippets relevant to functionalities based on natural language queries. Despite significant progress in self-supervised code pre-training utilizing the vast amount of code data in repositories, existing methods have primarily focused on leveraging contrastive learning to align natural language with function-level code snippets. These studies have overlooked the abundance of fine-grained (such as block-level and statement-level) code snippets prevalent within the function-level code snippets, which results in suboptimal performance across all levels of granularity. To address this problem, we first construct a multi-granularity code search dataset called MGCodeSearchNet, which contains 536K+ pairs of natural language and code snippets. Subsequently, we introduce a novel Multi-Granularity Self-Supervised contrastive learning code Search framework (MGS3). First, MGS3 features a Hierarchical Multi-Granularity Representation module (HMGR), which leverages syntactic structural relationships for hierarchical representation and aggregates fine-grained information into coarser-grained representations. Then, during the contrastive learning phase, we endeavor to construct positive samples of the same granularity for fine-grained code, and introduce in-function negative samples for fine-grained code. Finally, we conduct extensive experiments on code search benchmarks across various granularities, demonstrating that the framework exhibits outstanding performance in code search tasks of multiple granularities. These experiments also showcase its model-agnostic nature and compatibility with existing pre-trained code representation models.
Rui Li 0093, Junfeng Kang, Qi Liu 0003, Liyang He, Zheng Zhang 0048, Yunhao Sha, Linbo Zhu, Zhenya Huang
KDD (1)1
2024 CONSIDER: Commonalities and Specialties Driven Multilingual Code Retrieval Framework
abstract
Multilingual code retrieval aims to find code snippets relevant to a user's query from a multilingual codebase, which plays a crucial role in software development and expands their application scenarios compared to classical monolingual code retrieval. Despite the performance improvements achieved by previous studies, two crucial problems are overlooked in the multilingual scenario. First, certain programming languages face data scarcity in specific domains, resulting in limited representation capabilities within those domains. Second, different programming languages can be used interchangeably within the same domain, making it challenging for multilingual models to accurately identify the intended programming language of a user's query. To address these issues, we propose the CommONalities and SpecIalties Driven Multilingual CodE Retrieval Framework (CONSIDER), which includes two modules. The first module enhances the representation of various programming languages by modeling pairwise and global commonalities among them. The second module introduces a novel contrastive learning negative sampling algorithm that leverages language confusion to automatically extract specific language features. Through our experiments, we confirm the significant benefits of our model in real-world multilingual code retrieval scenarios in various aspects. Furthermore, an evaluation demonstrates the effectiveness of our proposed CONSIDER framework in monolingual scenarios as well. Our source code is available at https://github.com/smsquirrel/consider.
Rui Li 0093, Liyang He, Qi Liu 0003, Yuze Zhao, Zheng Zhang 0048, Zhenya Huang, Yu Su 0002, Shijin Wang 0001
AAAI1
2024 Reformulating Sequential Recommendation: Learning Dynamic User Interest with Content-enriched Language Modeling
Junzhe Jiang 0001, Shang Qu, Mingyue Cheng 0004, Qi Liu 0003, Zhiding Liu, Hao Zhang 0088, Rujiao Zhang, Kai Zhang 0038, Rui Li 0093, Jiatong Li 0002, Min Gao 0017
DASFAA (3)9
2024 Learning Recommender Systems with Soft Target: A Decoupled Perspective
Hao Zhang 0088, Mingyue Cheng 0004, Qi Liu 0003, Yucong Luo, Rui Li 0093, Enhong Chen
DASFAA (3)5
2024 Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models
abstract
Code retrieval aims to identify code from extensive codebases that semantically aligns with a given query code snippet.Collecting a broad and high-quality set of query and code pairs is crucial to the success of this task.However, existing data collection methods struggle to effectively balance scalability and annotation quality.In this paper, we first analyze the factors influencing the quality of function annotations generated by Large Language Models (LLMs).We find that the invocation of intra-repository functions and third-party APIs plays a significant role.Building on this insight, we propose a novel annotation method that enhances the annotation context by incorporating the content of functions called within the repository and information on third-party API functionalities.Additionally, we integrate LLMs with a novel sorting method to address the multi-level function call relationships within repositories.Furthermore, by applying our proposed method across a range of repositories, we have developed the Query4Code dataset.The quality of this synthesized dataset is validated through both model training and human evaluation, demonstrating high-quality annotations.Moreover, cost analysis confirms the scalability of our annotation method. 1
Rui Li 0093, Qi Liu 0003, Liyang He, Zheng Zhang 0048, Hao Zhang 0088, Shengyu Ye, Junyu Lu 0003, Zhenya Huang
EMNLP1
2024 A Unified Adaptive Testing System Enabled by Hierarchical Structure Search
abstract
Adaptive Testing System (ATS) is a promising testing mode, extensively utilized in standardized tests like the GRE. It offers personalized ability assessment by dynamically adjusting questions based on individual ability levels. Compared to traditional exams, ATS can improve the accuracy of ability estimates while simultaneously reducing the number of questions required. Despite the diverse testing formats of ATS, tailored to different adaptability requirements in various testing scenarios, there is a notable absence of a unified framework for modeling them. In this paper, we introduce a unified data-driven ATS framework that conceptualizes the various testing formats as a hierarchical test structure search problem. It can learn directly from data to solve for the optimal questions for each student, eliminating the need for manual test design. The proposed solution algorithm comes with theoretical guarantees for estimation error and convergence. Empirical results show that our framework maintains assessment accuracy while reducing question count by 20% on average and improving training stability.
Junhao Yu, Yan Zhuang 0001, Zhenya Huang, Qi Liu 0003, Xin Li 0064, Rui Li 0093, Enhong Chen
ICML6
2024 One-bit Deep Hashing: Towards Resource-Efficient Hashing Model with Binary Neural Network
abstract
Deep Hashing (DH) has emerged as an indispensable technique for fast image search in recent years. To deploy DH on resource-limited devices, the Binary Neural Network (BNN) offers a solution that significantly reduces computations and parameters compared to CNN. Unfortunately, applying BNN directly to DH will lead to huge performance degradation. To tackle this problem, we first conducted extensive experiments and discovered that the center-based method provides a fundamental guarantee for BNN-DH performance. Subsequently, we delved deeper into the impact of BNNs on center-based methods and revealed two key insights. First, we find reducing the distance between hash codes and hash centers is challenging for BNN-DH compared to CNN-based DH. Second, the evolution of hash code aggregation undergoes two stages in BNN-DH, which is different from CNN-based DH. Based on these findings, we designed a strong and general method called One-bit Deep Hashing (ODH). First, ODH incorporates a semantic self-adaptive hash center module to address the problem of hash codes inadequately converging to their hash centers. Then, it employs a novel two-stage training method to consider the evolution of hash code aggregation. Finally, extensive experiments on two datasets demonstrate that ODH can achieve significant superiority over other BNN-DH models.
Liyang He, Zhenya Huang, Rui Li 0093, Runze Wu 0001, Qi Liu 0003, Enhong Chen
ACM Multimedia4
2022 Co-promotion Predictions of Financing Market and Sales Market: A Cooperative-Competitive Attention Approach
abstract
Market popularity prediction has always been a hot research topic, such as sales prediction and crowdfunding prediction. Most of these studies put the perspective on isolated markets, relying on the knowledge of certain market to maximize the prediction performance. However, these market-specific approaches are restricted by the knowledge limitation of isolated markets and incapable of the complicated and potential relations among different markets, especially some with strong dependence such as the financing market and sales market. Fortunately, we discover potentially symbiotic relations between the financing market and the sales market, which provides us with an opportunity to co-promote the popularity predictions of both markets. Thus, for bridgly learning the knowledge interactions between financing market and sales market, we propose a cross-market approach, namely CATN: Cooperative-competitive Attention Transfer Network, which could effectively transfer knowledge of financing capability from the crowdfunding market and sales prospect from the E-commerce market. Specifically, for capturing the complicated relations especially the cooperation or complement of items and enhancing the knowledge transfer between the two heterogeneous markets, we design a novel Cooperative Attention; meanwhile, for finely computing the relations of items especially the competition in specific same market, we further design Competitive Attentions for the two markets respectively. Besides, we also distinguish aligned features and unique features to adapt the cross-market predictions. With the real-world datasets collected from Indiegogo and Amazon, we construct extensive experiments on three types of datasets from the two markets and the results demonstrate the effectiveness and generalization of our CATN model.
Lei Zhang 0060, Wang Xiang, Chuang Zhao 0002, Hongke Zhao, Rui Li 0093, Runze Wu 0001
AAAI5
2020 DLEP: A Deep Learning Model for Earthquake Prediction
abstract
Earthquakes are one of the most costly natural disasters facing human beings, which happens without an explicit warning, therefore earthquake prediction becomes a very important and challenging task for humanity. Although many existing methods attempt to address this task, most of them use either seismic indicators (explicit features) designed by geologists, or feature vectors (implicit features) extracted by deep learning methods, to characterize an earthquake for earthquake prediction. The problem of combining these two kind of features to improve final earthquake prediction performance remains pretty much open. To this end, we propose a deep learning model named DLEP to effectively fuse the explicit and implicit features for accurate earthquake prediction. In DLEP, we adopt eight precursory pattern-based indicators as the explicit features, and use a convolutional neural network (CNN) to extract implicit features. Then, an attention-based strategy is suggested to fuse these two kinds of features well. In addition, a dynamic loss function is designed to deal with the category imbalance of seismic data. Finally, experimental results on eight datasets from different regions demonstrate the effectiveness of the proposed DLEP for earthquake prediction comparing to several state-of-the-art baselines.
Rui Li 0093, Xiaobo Lu, Shuowei Li, Haipeng Yang, Jianfeng Qiu, Lei Zhang 0060
IJCNN1