Ruifeng Yuan

dblp:23/6662 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
abstract
Test-time compute scaling has emerged as a powerful paradigm for enhancing mathematical reasoning in large language models (LLMs) by allocating additional computational resources during inference. However, current methods employ uniform resource distribution across all reasoning sub-problems, creating fundamental bottlenecks where challenging sub-problems receive insufficient attention while routine operations consume disproportionate resources. This uniform allocation creates performance bottlenecks where additional computational resources yield diminishing returns. Inspired by dual-process theory, we propose SCALE (Selective Resource Allocation), a framework that selectively allocates computational resources based on sub-problem difficulty. SCALE operates through four stages: (1) problem decomposition into sequential reasoning sub-problems, (2) difficulty assessment of each sub-problem to distinguish between routine operations and computationally challenging sub-problems, (3) selective processing mode assignment between System 1 for simple sub-problems and System 2 for complex ones, and (4) sequential execution with context propagation. By concentrating resources on challenging sub-problems while processing routine operations efficiently, SCALE achieves substantial performance improvements with superior resource utilization. Extensive experiments demonstrate that SCALE significantly outperforms uniform scaling baselines, achieving accuracy improvements of up to 13.75 percentage points (57.50% to 71.25% on AIME25) while reducing computational costs by 33-53%, representing a major advance in test-time scaling that addresses fundamental limitations of current approaches.
Chunpu Xu, Ruifeng Yuan, Jessie Wang 0004, Wenjie Li 0002, Pengfei Liu 0003
AAAI3
2026 CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation
abstract
Ruifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi, Zhongyu Wei, Ling Zhang, Jianpeng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruifeng Yuan, Wanxing Chang, Zhongyu Wei, Ling Zhang 0002
ACL (1)1
2026 Understanding the Behaviors of Environment-aware Information Retrieval
abstract
Ruifeng Yuan, Chaohao Yuan, David Dai, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruifeng Yuan, Chaohao Yuan, David Dai, Yu Rong 0001, Hong Cheng 0001, Hou Pong Chan, Chenghao Xiao
ACL (1)1
2025 OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
abstract
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Jian Yang 0030, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang 0001, Yang Gao 0021, Jie Fu 0001, Qian Liu 0033, Houyi Li, Ge Zhang 0009, Yuan Qi 0001
ACL (1)12
2025 Personalized Large Language Model Assistant with Evolving Conditional Memory
abstract
With the rapid development of large language models, AI assistants like ChatGPT have become increasingly integrated into people’s works and lives but are limited in personalized services. In this paper, we present a plug-and-play framework that could facilitate personalized large language model assistants with evolving conditional memory. The personalized assistant focuses on intelligently preserving the knowledge and experience from the history dialogue with the user, which can be applied to future tailored responses that better align with the user’s preferences. Generally, the assistant generates a set of records from the dialogue, stores them in a memory bank, and retrieves related memory to improve the quality of the response. For the crucial memory design, we explore different ways of constructing the memory and propose a new memorizing mechanism named conditional memory to enhance the memory management of the framework. We also investigate the retrieval and usage of memory in the generation process. To better evaluate the personalized assistants’ abilities, we build the first evaluation benchmark from three critical aspects: continuing previous dialogue, learning personalized knowledge and learning from user feedback. The experimental results illustrate the effectiveness of our method.
Ruifeng Yuan, Shichao Sun, Yongqi Li 0001, Ziqiang Cao, Wenjie Li 0002
COLING1
2025 LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
abstract
Large language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose elements that mirror human problem-solving, categorized as progressive reasoning (the essential solution development path) and functional elements (verification processes, alternative solution approaches, and error corrections). While progressive reasoning is crucial, the functional elements significantly increase computational demands during test-time inference. We introduce PIR (Perplexity-based Importance Refinement), a principled framework that quantitatively evaluates the importance of each reasoning step based on its impact on answer prediction confidence. PIR systematically identifies and selectively prunes only low-importance functional steps while preserving all progressive reasoning components, creating optimized training data that maintains the integrity of the core solution path while reducing verbosity. Models fine-tuned on PIR-optimized data exhibit superior test-time scaling properties, generating more concise reasoning chains while achieving improved accuracy (+0.9\% to +6.6\%) with significantly reduced token usage (-3\% to -41\%) across challenging reasoning benchmarks (AIME, AMC, and GPQA Diamond). Our approach demonstrates strong generalizability across different model sizes, data sources, and token budgets, offering a practical solution for deploying reasoning-capable LLMs in scenarios where efficient test-time scaling, response time, and computational efficiency are valuable constraints. Code and dataset are available at the [LIMOPro GitHub repository.](https://github.com/GAIR-NLP/LIMOPro)
Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li 0002, Pengfei Liu 0003
NeurIPS3
2025 Exploring Training and Inference Scaling Laws in Generative Retrieval
abstract
Generative retrieval reformulates retrieval as an autoregressive generation task, where large language models (LLMs) generate target documents directly from a query. As a novel paradigm, the mechanisms that underpin its performance and scalability remain largely unexplored. We systematically investigate training and inference scaling laws in generative retrieval, exploring how model size, training data scale, and inference-time compute jointly influence performance. We propose a novel evaluation metric inspired by contrastive entropy and generation loss, providing a continuous performance signal that enables robust comparisons across diverse generative retrieval methods. Our experiments show that n-gram-based methods align strongly with training and inference scaling laws. We find that increasing model size, training data scale, and inference-time compute all contribute to improved performance, highlighting the complementary roles of these factors in enhancing generative retrieval. Across these settings, LLaMA models consistently outperform T5 models, suggesting a particular advantage for larger decoder-only models in generative retrieval. Our findings underscore that model sizes, data availability, and inference computation interact to unlock the full potential of generative retrieval, offering new insights for designing and optimizing future systems. We release code at SLGR GitHub repository.
Hongru Cai, Yongqi Li 0001, Ruifeng Yuan, Wenjie Wang 0007, Zhen Zhang 0008, Wenjie Li 0002, Tat-Seng Chua
SIGIR3
2024 QuerySum: A Multi-Document Query-Focused Summarization Dataset Augmented with Similar Query Clusters
abstract
Query-focused summarization (QFS) aims to summarize the source document(s) with regard to a specific aspect of information given in a query. It plays an important role in presenting users with a concise answer summary from a set of query-relevant documents retrieved by the information retrieval system. Nonetheless, the QFS research has long been hampered by the lack of adequate datasets in terms of both quality and quantity. In this paper, we introduce a large-scale multi-document query-focused summarization dataset, called QuerySum, which contains 27,041 data samples covering diverse topics and its quality is guaranteed through human verification. Unlike some previous QFS datasets constructed directly from the question answering datasets, 74% queries in our dataset are the challenging non-factoid What-, Why-, and How- questions. More importantly, we also provide a set of similar queries together with the corresponding summaries pairs for each query as the retrieved context, presenting a new feature of QuerySum. We aim to encourage research efforts in query intention understanding in the context of QFS. Leveraging QuerySum's depth, we propose a model for query-aware multi-document summarization and set a new QFS benchmark.
Ruifeng Yuan
AAAI3
2024 Dr. CLIP: CLIP-Driven Universal Framework for Zero-Shot Sketch Image Retrieval
abstract
The field of Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is undergoing a paradigm shift, transitioning from specialized models designed for individual tasks to more general retrieval models capable of managing various specialized scenarios. Inspired by the impressive generalization ability of the Contrastive Language-Image Pretraining (CLIP) model, we propose a CLIP-driven universal framework (Dr. CLIP), which leverages prompt learning to guide the synergy between CLIP and ZS-SBIR. Dr. CLIP can perfectly cover four variants of ZS-SBIR tasks (inter-category, intra-category, cross-datasets, and generalization). Moreover, we decompose the synergy into classification learning, metric learning, and ranking learning, as well as introduce three key components to enhance learning effectiveness. i) a forgetting suppression idea is applied to prevent catastrophic forgetting and constrains the feature distribution of the new categories in classification learning. ii) a domain balanced loss is proposed to address sample imbalance and establish effective cross-domain correlations in metric learning. iii) a pair-relation strategy is introduced to capture relevance and ranking relationships between instances in ranking learning. Eventually, we reorganize and redivide three coarse-grained datasets and two fine-grained datasets to accommodate the training settings for four ZS-SBIR tasks. The comparison experiments confirmed our method surpassed the state-of-the-art (SOTA) methods by a significant margin (1.95% ~ 19.14,% mAP), highlighting its generality and superiority. The code is available at https://github.com/x-28/CDUF.git.
Xue Li 0008, Ziyang Li 0010, Hongchun Lu, Ruifeng Yuan
ACM Multimedia5
2024 Dialogue acts enhanced extract-abstract framework for meeting summarization
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Ziqiang Cao, Sujian Li
Inf. Process. Manag.2
2023 Preserve Context Information for Extract-Generate Long-Input Summarization Framework
abstract
The Extract-generate framework has been a classic approach for text summarization. As pretrained language models struggling with long-input summarization for their high memory cost, extract-generate framework regains researchers' interests. However, the cost of its effectiveness in dealing with long-input summarization is the loss of context information. In this paper, we present a context-aware extract-generate framework (CAEG) for long-input text summarization. It focuses on preserving both local and global context information in an extract-generate framework with little cost, and can be applied to most of existing extract-generate summarization models. CAEG generates a set of context-related text spans called context prompts for each text snippet and use them to transfer the context information from the extractor and generator. To find such context prompts, we propose to capture the context information based on the interpretation of the extractor, where the text spans having the highest contribution to the extraction decision are considered as containing the richest context information. We evaluate our approach on both long-document and long-dialogue summarization datasets: arXiv and QMSum. The experiment results show that CAEG achieves the-state-of-art result on QMSum and outperforms other extract-generate based models in arXiv.
Ruifeng Yuan, Ziqiang Cao, Wenjie Li 0002
AAAI1
2023 Improving Sentence Similarity Estimation for Unsupervised Extractive Summarization
abstract
Unsupervised extractive summarization aims to extract salient sentences from a document as the summary without labeled data. Recent literatures mostly research how to leverage sentence similarity to rank sentences in the order of salience. However, sentence similarity estimation using pre-trained language models mostly takes little account of document-level information and has a weak correlation with sentence salience ranking. In this paper, we proposed two novel strategies to improve sentence similarity estimation for unsupervised extractive summarization. We use contrastive learning to optimize a document-level objective that sentences from the same document are more similar than those from different documents. Moreover, we use mutual learning to enhance the relationship between sentence similarity estimation and sentence salience ranking, where an extra signal amplifier is used to refine the pivotal information. Experimental results demonstrate the effectiveness of our strategies.1
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Sujian Li
ICASSP2
2022 Few-shot Query-Focused Summarization with Prefix-Merging
abstract
Query-focused summarization has been considered as an important extension for text summarization.It aims to generate a concise highlight for a given query.Different from text summarization, query-focused summarization has long been plagued by the problem of lacking high-quality large-scale datasets.In this paper, we investigate the idea that whether we can integrate and transfer the knowledge of text summarization and question answering to assist the few-shot learning in query-focused summarization.Here, we propose prefix-merging, a prefix-based pretraining strategy for fewshot learning in query-focused summarization.Drawn inspiration from prefix-tuning, we are allowed to integrate the task knowledge from text summarization and question answering into a properly designed prefix and apply the merged prefix to query-focused summarization.With only a small amount of trainable parameters, prefix-merging outperforms fine-tuning on query-focused summarization.We further discuss the influence of different prefix designs and propose a visualized explanation for how prefix-merging works.
Ruifeng Yuan, Ziqiang Cao, Wenjie Li 0002
EMNLP1
2021 Event Graph based Sentence Fusion
abstract
Sentence fusion is a conditional generation task that merges several related sentences into a coherent one, which can be deemed as a summary sentence.The importance of sentence fusion has long been recognized by communities in natural language generation, especially in text summarization.It remains challenging for a state-of-the-art neural abstractive summarization model to generate a well-integrated summary sentence.In this paper, we explore the effective sentence fusion method in the context of text summarization.We propose to build an event graph from the input sentences to effectively capture and organize related events in a structured way and use the constructed event graph to guide sentence fusion.In addition to make use of the attention over the content of sentences and graph nodes, we further develop a graph flow attention mechanism to control the fusion process via the graph structure.When evaluated on sentence fusion data built from two summarization datasets, CNN/DaliyMail and Multi-News, our model shows to achieve state-of-the-art performance in terms of Rouge and other metrics like fusion rate and faithfulness.
Ruifeng Yuan
EMNLP (1)1
2020 Fact-level Extractive Summarization with Hierarchical Graph Mask on BERT
abstract
Most current extractive summarization models generate summaries by selecting salient sentences.However, one of the problems with sentence-level extractive summarization is that there exists a gap between the human-written gold summary and the oracle sentence labels.In this paper, we propose to extract fact-level semantic units for better extractive summarization.We also introduce a hierarchical structure, which incorporates the multi-level of granularities of the textual information into the model.In addition, we incorporate our model with BERT using a hierarchical graph mask.This allows us to combine BERT's ability in natural language understanding and the structural information without increasing the scale of the model.Experiments on the CNN/DaliyMail dataset show that our model achieves state-of-the-art results.
Ruifeng Yuan
COLING1
2019 dTexSL: A dynamic disaster textual storyline generating framework
Ruifeng Yuan, Qifeng Zhou, Wubai Zhou
World Wide Web1