Yifan Gao 0001

dblp:79/3190-1 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-1881-4577ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Leveraging historical information to boost retrieval-augmented generation in conversations
abstract
Multi-turn interactions between users and information-seeking systems have become a popular paradigm to satisfy complex information needs via a flexible interface and context understanding capacity. However, existing methods primarily adapt single-turn retrieval-augmented generation (RAG) pipelines to conversational settings without effectively incorporating historical information, such as previous search results, turn dependency, and historical evidence grounding. To effectively manage and utilize the information in conversations, we explore the feasibility of boosting response generation by leveraging historical information and propose several strategies to incorporate this information individually or in combination. We conduct experiments on three widely used conversational search benchmarks, each containing thousands of samples. Our method consistently outperforms previous strong baselines across different settings, achieving approximately a 10% absolute improvement over the second-best approach. Besides, our analyses help to understand the behind-the-scenes behavior of our methods. • We investigate the feasibility of leveraging abundant historical information to improve RAG performance in conversations. • We design several training-free strategies from different aspects, that can be used individually or in combination to boost RAG performance. • We conduct thorough experiments on three datasets to demonstrate the effectiveness of our methods, and analyze the potential paradigms behind the model.
Fengran Mo, Yifan Gao 0001, Zhuofeng Wu 0005, Xin Liu 0039, Zheng Li 0018, Meng Jiang 0001, Jian-Yun Nie
Inf. Process. Manag.2
2025 UniConv: Unifying Retrieval and Response Generation for Large Language Models in Conversations
abstract
Fengran Mo, Yifan Gao, Chuan Meng, Xin Liu, Zhuofeng Wu, Kelong Mao, Zhengyang Wang, Pei Chen, Zheng Li, Xian Li, Bing Yin, Meng Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fengran Mo, Yifan Gao 0001, Chuan Meng, Xin Liu 0039, Zhuofeng Wu 0005, Kelong Mao, Zheng Li 0018, Meng Jiang 0001
ACL (1)2
2025 Aligning Large Language Models with Implicit Preferences from User-Generated Content
abstract
Learning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from humans or advanced LLMs, which is costly and difficult to scale. In this work, we present PUGC, a novel framework that leverages implicit human Preferences in unlabeled User-Generated Content (UGC) to generate preference data. Although UGC is not explicitly created to guide LLMs in generating human-preferred responses, it often reflects valuable insights and implicit preferences from its creators that has the potential to address readers’ questions. PUGC transforms UGC into user queries and generates responses from the policy model. The UGC is then leveraged as a reference text for response scoring, aligning the model with these implicit preferences. This approach improves the quality of preference data while enabling scalable, domain-specific alignment. Experimental results on Alpaca Eval 2 show that models trained with DPO and PUGC achieve a 9.37% performance improvement over traditional methods, setting a 35.93% state-of-the-art length-controlled win rate using Mistral-7B-Instruct. Further studies highlight gains in reward quality, domain-specific alignment effectiveness, robustness against UGC quality, and theory of mind capabilities. Our code and dataset are available at https://zhaoxuan.info/PUGC.github.io/.
Zhaoxuan Tan, Zheng Li 0018, Hyokun Yun, Ming Zeng 0001, Zhihan Zhang 0001, Yifan Gao 0001, Ruijie Wang 0004, Priyanka Nigam, Meng Jiang 0001
ACL (1)9
2025 EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association
abstract
Weiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag, Wenju Xu, Chen Luo, Sheikh Muhammad Sarwar, Yang Li, Hansu Gu, Hui Liu, Changlong Yu, Jiaxin Bai, Yifan Gao, Haiyang Zhang, Qi He, Shuiwang Ji, Yangqiu Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weiqi Wang 0001, Limeng Cui, Xin Liu 0039, Sreyashi Nag, Wenju Xu, Chen Luo 0003, Sheikh Muhammad Sarwar, Yang Li 0055, Hansu Gu, Hui Liu 0033, Changlong Yu, Jiaxin Bai, Yifan Gao 0001, Qi He 0002, Shuiwang Ji, Yangqiu Song
ACL (1)13
2025 M+: Extending MemoryLLM with Scalable Long-Term Memory
abstract
Equipping large language models (LLMs) with latent-space memory has attracted increasing attention as they can extend the context window of existing language models. However, retaining information from the distant past remains a challenge. For example, MemoryLLM (Wang et al., 2024a), as a representative work with latent-space memory, compresses past information into hidden states across all layers, forming a memory pool of 1B parameters. While effective for sequence lengths up to 16k tokens, it struggles to retain knowledge beyond 20k tokens. In this work, we address this limitation by introducing M+, a memory-augmented model based on MemoryLLM that significantly enhances long-term information retention. M+ integrates a long-term memory mechanism with a co-trained retriever, dynamically retrieving relevant information during text generation. We evaluate M+ on diverse benchmarks, including long-context understanding and knowledge retention tasks. Experimental results show that M+ significantly outperforms MemoryLLM and recent strong baselines, extending knowledge retention from under 20k to over 160k tokens with similar GPU memory overhead.
Yu Wang 0170, Dmitry Krotov, Yifan Gao 0001, Wangchunshu Zhou, Julian J. McAuley, Dan Gutfreund, Rogério Feris, Zexue He
ICML4
2025 ALERT: An LLM-powered Benchmark for Automatic Evaluation of Recommendation Explanations
abstract
Yichuan Li, Xinyang Zhang, Chenwei Zhang, Mao Li, Tianyi Liu, Pei Chen, Yifan Gao, Kyumin Lee, Kaize Ding, Zhengyang Wang, Zhihan Zhang, Jingbo Shang, Xian Li, Trishul Chilimbi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yichuan Li 0001, Xinyang Zhang 0002, Yifan Gao 0001, Kyumin Lee, Kaize Ding, Zhihan Zhang 0001, Jingbo Shang, Trishul Chilimbi
NAACL (Long Papers)7
2025 IHEval: Evaluating Language Models on Following the Instruction Hierarchy
abstract
Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu, Haoming Jiang, Xianfeng Tang, Yifan Gao, Zheng Li, Haodong Wang, Zhaoxuan Tan, Yichuan Li, Qingyu Yin, Bing Yin, Meng Jiang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhihan Zhang 0001, Xin Liu 0039, Haoming Jiang, Xianfeng Tang, Yifan Gao 0001, Zheng Li 0018, Zhaoxuan Tan, Yichuan Li 0001, Qingyu Yin, Meng Jiang 0001
NAACL (Long Papers)7
2025 Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
abstract
Yuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao, Qing Ping, Tianyi Liu, Binxuan Huang, Zheng Li, Zhengyang Wang, Pei Chen, Ruijie Wang, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Bing Yin, Chao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yuchen Zhuang, Jingfeng Yang 0001, Haoming Jiang, Xin Liu 0039, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao 0001, Qing Ping, Binxuan Huang, Zheng Li 0018, Ruijie Wang 0004, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Chao Zhang 0014
NAACL (Long Papers)7
2024 Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark
abstract
Fenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng, Haoming Jiang, Yifan Gao, Priyanka Nigam, Sreyashi Nag, Bing Yin, Yining Hua, Xuan Zhou, Omid Rohanian, Anshul Thakur, Lei Clifton, David A. Clifton. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zheng Li 0018, Hongjian Zhou, Qingyu Yin, Jingfeng Yang 0001, Xianfeng Tang, Chen Luo 0003, Ming Zeng 0001, Haoming Jiang, Yifan Gao 0001, Priyanka Nigam, Sreyashi Nag, Yining Hua, Omid Rohanian, Anshul Thakur, Lei A. Clifton, David A. Clifton
EMNLP10
2024 MEMORYLLM: Towards Self-Updatable Large Language Models
abstract
Existing Large Language Models (LLMs) usually remain static after deployment, which might make it hard to inject new knowledge into the model. We aim to build models containing a considerable portion of self-updatable parameters, enabling the model to integrate new knowledge effectively and efficiently. To this end, we introduce MEMORYLLM, a model that comprises a transformer and a fixed-size memory pool within the latent space of the transformer. MEMORYLLM can self-update with text knowledge and memorize the knowledge injected earlier. Our evaluations demonstrate the ability of MEMORYLLM to effectively incorporate new knowledge, as evidenced by its performance on model editing benchmarks. Meanwhile, the model exhibits long-term information retention capacity, which is validated through our custom-designed evaluations and long-context benchmarks. MEMORYLLM also shows operational integrity without any sign of performance degradation even after nearly a million memory updates. Our code and model are open-sourced at https://github.com/wangyu-ustc/MemoryLLM.
Yu Wang 0170, Yifan Gao 0001, Xiusi Chen, Haoming Jiang, Jingfeng Yang 0001, Qingyu Yin, Zheng Li 0018, Jingbo Shang, Julian J. McAuley
ICML2
2024 Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language Models
abstract
Online shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks. However, existing models and benchmarks are commonly tailored to specific tasks, falling short of capturing the full complexity of online shopping. Large Language Models (LLMs), with their multi-task and few-shot learning abilities, have the potential to profoundly transform online shopping by alleviating task-specific engineering efforts and by providing users with interactive conversations. Despite the potential, LLMs face unique challenges in online shopping, such as domain-specific concepts, implicit knowledge, and heterogeneous user behaviors. Motivated by the potential and challenges, we propose Shopping MMLU, a diverse multi-task online shopping benchmark derived from real-world Amazon data. Shopping MMLU consists of 57 tasks covering 4 major shopping skills: concept understanding, knowledge reasoning, user behavior alignment, and multi-linguality, and can thus comprehensively evaluate the abilities of LLMs as general shop assistants. With Shoppping MMLU, we benchmark over 20 existing LLMs and uncover valuable insights about practices and prospects of building versatile LLM-based shop assistants. Shopping MMLU can be publicly accessed at https://github.com/KL4805/ShoppingMMLU. In addition, with Shopping MMLU, we are hosting a competition in KDD Cup 2024 with over 500 participating teams. The winning solutions and the associated workshop can be accessed at our website https://amazon-kddcup24.github.io/.
Yilun Jin, Zheng Li 0018, Tianyu Cao 0001, Yifan Gao 0001, Pratik Jayarao, Xin Liu 0039, Ritesh Sarkhel, Xianfeng Tang, Wenju Xu, Jingfeng Yang 0001, Qingyu Yin, Priyanka Nigam, Yi Xu 0011, Kai Chen 0005, Qiang Yang 0001, Meng Jiang 0001
NeurIPS5
2023 SCOTT: Self-Consistent Chain-of-Thought Distillation
abstract
Large language models (LMs) beyond a certain scale, demonstrate the emergent capability of generating free-text rationales for their predictions via chain-of-thought (CoT) prompting.While CoT can yield dramatically improved performance, such gains are only observed for sufficiently large LMs.Even more concerning, there is little guarantee that the generated rationales are consistent with LM's predictions or faithfully justify the decisions.In this work, we propose SCOTT, a faithful knowledge distillation method to learn a small, self-consistent CoT model from a teacher model that is orders of magnitude larger.To form better supervision, we elicit rationales supporting the gold answers from a large LM (teacher) by contrastive decoding, which encourages the teacher to generate tokens that become more plausible only when the answer is considered.To ensure faithful distillation, we use the teacher-generated rationales to learn a student LM with a counterfactual reasoning objective, which prevents the student from ignoring the rationales to make inconsistent predictions.Experiments show that, while yielding comparable end-task performance, our method can generate CoT rationales that are more faithful than baselines do.Further analysis suggests that such a model respects the rationales more when making decisions; thus, we can improve its performance more by refining its rationales.
Peifeng Wang, Zheng Li 0018, Yifan Gao 0001, Xiang Ren 0001
ACL (1)4
2023 Enhancing User Intent Capture in Session-Based Recommendation with Attribute Patterns
abstract
The goal of session-based recommendation in E-commerce is to predict the next item that an anonymous user will purchase based on the browsing and purchase history. However, constructing global or local transition graphs to supplement session data can lead to noisy correlations and user intent vanishing. In this work, we propose the Frequent Attribute Pattern Augmented Transformer (FAPAT) that characterizes user intents by building attribute transition graphs and matching attribute patterns. Specifically, the frequent and compact attribute patterns are served as memory to augment session representations, followed by a gate and a transformer block to fuse the whole session information. Through extensive experiments on two public benchmarks and 100 million industrial data in three domains, we demonstrate that FAPAT consistently outperforms state-of-the-art methods by an average of 4.5% across various evaluation metrics (Hits, NDCG, MRR). Besides evaluating the next-item prediction, we estimate the models' capabilities to capture user intents via predicting items' attributes and period-item recommendations.
Xin Liu 0039, Zheng Li 0018, Yifan Gao 0001, Jingfeng Yang 0001, Tianyu Cao 0001, Yangqiu Song
NeurIPS3
2022 ProQA: Structural Prompt-based Pre-training for Unified Question Answering
abstract
Wanjun Zhong, Yifan Gao, Ning Ding, Yujia Qin, Zhiyuan Liu, Ming Zhou, Jiahai Wang, Jian Yin, Nan Duan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Wanjun Zhong, Yifan Gao 0001, Ning Ding 0002, Yujia Qin, Zhiyuan Liu 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001
NAACL-HLT2
2022 Query Attribute Recommendation at Amazon Search
abstract
Query understanding models extract attributes from search queries, like color, product type, brand, etc. Search engines rely on these attributes for ranking, advertising, and recommendation, etc. However, product search queries are usually short, three or four words on average. This information shortage limits the search engine’s power to provide high-quality services.
Chen Luo 0003, William Headden, Neela Avudaiappan, Haoming Jiang, Tianyu Cao 0001, Qingyu Yin, Yifan Gao 0001, Zheng Li 0018, Rahul Goutam
RecSys7
2021 Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction
abstract
Yifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)1
2020 Explicit Memory Tracker with Coarse-to-Fine Reasoning for Conversational Machine Reading
abstract
Yifan Gao, Chien-Sheng Wu, Shafiq Joty, Caiming Xiong, Richard Socher, Irwin King, Michael Lyu, Steven C.H. Hoi. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Yifan Gao 0001, Chien-Sheng Wu, Shafiq R. Joty, Caiming Xiong, Richard Socher, Irwin King, Michael R. Lyu, Steven C. H. Hoi
ACL1
2020 Discern: Discourse-Aware Entailment Reasoning Network for Conversational Machine Reading
abstract
Yifan Gao, Chien-Sheng Wu, Jingjing Li, Shafiq Joty, Steven C.H. Hoi, Caiming Xiong, Irwin King, Michael Lyu. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Yifan Gao 0001, Chien-Sheng Wu, Jingjing Li 0007, Shafiq R. Joty, Steven C. H. Hoi, Caiming Xiong, Irwin King, Michael R. Lyu
EMNLP (1)1
2019 Title-Guided Encoding for Keyphrase Generation
abstract
Keyphrase generation (KG) aims to generate a set of keyphrases given a document, which is a fundamental task in natural language processing (NLP). Most previous methods solve this problem in an extractive manner, while recently, several attempts are made under the generative setting using deep neural networks. However, the state-of-the-art generative methods simply treat the document title and the document main body equally, ignoring the leading role of the title to the overall document. To solve this problem, we introduce a new model called Title-Guided Network (TG-Net) for automatic keyphrase generation task based on the encoderdecoder architecture with two new features: (i) the title is additionally employed as a query-like input, and (ii) a titleguided encoder gathers the relevant information from the title to each word in the document. Experiments on a range of KG datasets demonstrate that our model outperforms the state-of-the-art models with a large margin, especially for documents with either very low or very high title length ratios.
Wang Chen 0001, Yifan Gao 0001, Jiani Zhang 0001, Irwin King, Michael R. Lyu
AAAI2
2019 Generating Distractors for Reading Comprehension Questions from Real Examinations
abstract
We investigate the task of distractor generation for multiple choice reading comprehension questions from examinations. In contrast to all previous works, we do not aim at preparing words or short phrases distractors, instead, we endeavor to generate longer and semantic-rich distractors which are closer to distractors in real reading comprehension from examinations. Taking a reading comprehension article, a pair of question and its correct option as input, our goal is to generate several distractors which are somehow related to the answer, consistent with the semantic context of the question and have some trace in the article. We propose a hierarchical encoderdecoder framework with static and dynamic attention mechanisms to tackle this task. Specifically, the dynamic attention can combine sentence-level and word-level attention varying at each recurrent time step to generate a more readable sequence. The static attention is to modulate the dynamic attention not to focus on question irrelevant sentences or sentences which contribute to the correct option. Our proposed framework outperforms several strong baselines on the first prepared distractor generation dataset of real reading comprehension questions. For human evaluation, compared with those distractors generated by baselines, our generated distractors are more functional to confuse the annotators.
Yifan Gao 0001, Lidong Bing, Piji Li, Irwin King, Michael R. Lyu
AAAI1
2019 Interconnected Question Generation with Coreference Alignment and Conversation Flow Modeling
abstract
We study the problem of generating interconnected questions in question-answering style conversations.Compared with previous works which generate questions based on a single sentence (or paragraph), this setting is different in two major aspects: (1) Questions are highly conversational.Almost half of them refer back to conversation history using coreferences.(2) In a coherent conversation, questions have smooth transitions between turns.We propose an end-to-end neural model with coreference alignment and conversation flow modeling.The coreference alignment modeling explicitly aligns coreferent mentions in conversation history with corresponding pronominal references in generated questions, which makes generated questions interconnected to conversation history.The conversation flow modeling builds a coherent conversation by starting questioning on the first few sentences in a text passage and smoothly shifting the focus to later parts.Extensive experiments show that our system outperforms several baselines and can generate highly conversational questions.The code implementation is released at https://github.com/ Evan-Gao/conversaional-QG.
Yifan Gao 0001, Piji Li, Irwin King, Michael R. Lyu
ACL (1)1
2019 Improving Question Generation With to the Point Context
abstract
Jingjing Li, Yifan Gao, Lidong Bing, Irwin King, Michael R. Lyu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jingjing Li 0007, Yifan Gao 0001, Lidong Bing, Irwin King, Michael R. Lyu
EMNLP/IJCNLP (1)2
2019 Difficulty Controllable Generation of Reading Comprehension Questions
abstract
We investigate the difficulty levels of questions in reading comprehension datasets such as SQuAD, and propose a new question generation setting, named Difficulty-controllable Question Generation (DQG). Taking as input a sentence in the reading comprehension paragraph and some of its text fragments (i.e., answers) that we want to ask questions about, a DQG method needs to generate questions each of which has a given text fragment as its answer, and meanwhile the generation is under the control of specified difficulty labels---the output questions should satisfy the specified difficulty as much as possible. To solve this task, we propose an end-to-end framework to generate questions of designated difficulty levels by exploring a few important intuitions. For evaluation, we prepared the first dataset of reading comprehension questions with difficulty labels. The results show that the question generated by our framework not only have better quality under the metrics like BLEU, but also comply with the specified difficulty labels.
Yifan Gao 0001, Lidong Bing, Wang Chen 0001, Michael R. Lyu, Irwin King
IJCAI1
2017 Learning multi-level features for sensor-based human action recognition
Yan Xu 0001, Zhengyang Shen, Yifan Gao 0001, Shujian Deng, Yubo Fan, Eric I-Chao Chang
Pervasive Mob. Comput.4