Feiyun Ouyang

dblp:359/6097 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-7061-7351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 70% Multi-agent systems · 7% Knowledge representation and reasoning · 7%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 50% Human-AI interaction · 50%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
retrieval-augmented generation
2.732026
PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations · EMNLP 2025
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
knowledge-grounded reasoning
1.012026
PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026
Natural language and speech › Language models and text generation
large language model reasoning
1.012026
PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems
multi-agent reasoning
1.012026
PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
multi-step planning
1.012026
PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026
Health and well-being technologies › behavior change
health behavior change
1.012026
ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care · AAAI 2026
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality
0.912025
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering
0.912025
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations · EMNLP 2025
Natural language and speech › Language models and text generation › large language model reasoning
reasoning enhancement
0.912025
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation
alignment
0.812024
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024
Natural language and speech › Language models and text generation › text summarization › biomedical summarization
clinical summarization
0.812024
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
factual consistency
0.812024
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024
Natural language and speech › Language models and text generation
hallucination mitigation
0.812024
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024
Natural language and speech › Language models and text generation
text summarization
0.812024
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024
Medical and health informatics › digital health
diabetes management
0.312026
ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
0.312025
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations · EMNLP 2025
Natural language and speech › Language models and text generation
large language model
0.312025
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

large language model · 3.6persuasive strategy · 2.0multi-agent simulation · 2.0retrieval · 1.0multi-agent framework · 1.0chain-of-thought · 1.0sub-question decomposition · 0.9retrieval-augmented generation · 0.9retrieval augmentation · 0.9factuality scoring · 0.9code execution · 0.9DPO · 0.8
YearPublicationVenuePosition
2026 PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
abstract
Inspired by the dual-process theory of human cognition from Thinking, Fast and Slow, we introduce PRIME (Planning and Retrieval-Integrated Memory for Enhanced Reasoning), a multi-agent reasoning framework that dynamically integrates System 1 (fast, intuitive thinking) and System 2 (slow, deliberate thinking). PRIME first employs a Quick Thinking Agent to generate a rapid answer; if uncertainty is detected, it then triggers a structured System 2 reasoning pipeline composed of specialized agents for planning, hypothesis generation, retrieval, information integration, and decision-making. This multi-agent design mimics human cognitive processes faithfully and enhances both efficiency and accuracy. Experimental results with LLaMA 3 models demonstrate that PRIME enables open-source LLMs to perform competitively with state-of-the-art closed-source models like GPT-4 and GPT-4o on benchmarks requiring multi-hop and knowledge-grounded reasoning. This research establishes PRIME as a scalable solution for improving LLMs in domains requiring complex, knowledge-intensive reasoning.
Hieu Tran, Zonghai Yao, Nguyen Luong Tran, Zhichao Yang 0001, Feiyun Ouyang, Razieh Rahimi, Hong Yu 0001
AAAI5
2026 ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care
abstract
Real-world adoption of closed-loop insulin delivery systems (CLIDS) in type 1 diabetes remains low, driven not by technical failure, but by diverse behavioral, psychosocial, and social barriers. We introduce ChatCLIDS, the first benchmark to rigorously evaluate LLM–driven persuasive dialogue for health behavior change. Our framework features a library of expert-validated virtual patients, each with clinically grounded, heterogeneous profiles and realistic adoption barriers, and simulates multi-turn interactions with nurse agents equipped with a diverse set of evidence-based persuasive strategies. ChatCLIDS uniquely supports longitudinal counseling and adversarial social influence scenarios, enabling robust, multi-dimensional evaluation. Our findings reveal that while larger and more reflective LLMs adapt strategies over time, all models struggle to overcome resistance, especially under realistic social pressure. These results highlight critical limitations of current LLMs for behavior change, and offer a high-fidelity, scalable testbed for advancing trustworthy persuasive AI in healthcare and beyond.
Zonghai Yao, Talha Chafekar, Junda Wang, Feiyun Ouyang, Junhui Qian, Hong Yu 0001
AAAI5
2025 RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
abstract
, which leverages information retrieval specifically for generated sub-questions and re-answers these sub-questions with the relevant contextual information. Additionally, a Retrieval-Augmented Factuality Scorer is proposed to replace the original discriminator, prioritizing reasoning paths that meet high standards of factuality. Experimental results with LLaMA 3.1 show that RARE enables open-source LLMs to achieve competitive performance with top closed-source models like GPT-4 and GPT-4o. This research establishes RARE as a scalable solution for improving LLMs in domains where logical coherence and factual integrity are critical.
Hieu Tran, Zonghai Yao, Zhichao Yang 0001, Junda Wang, Feiyun Ouyang, Hong Yu 0001
ACL (1)7
2025 From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
abstract
Large language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical decision-making, remains underexplored and poorly evaluated. Existing benchmarks often assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. In this work, we revisit medical calculation evaluation with a stronger focus on clinical trustworthiness. First, we clean and restructure the MedCalc-Bench dataset and propose a new step-by-step evaluation pipeline that independently assesses formula selection, entity extraction, and arithmetic computation. Under this granular framework, the accuracy of GPT-4o drops from 62.7% to 43.6%, revealing errors masked by prior evaluations. Second, we introduce an automatic error analysis framework that generates structured attribution for each failure mode. Human evaluation confirms its alignment with expert judgment, enabling scalable and explainable diagnostics. Finally, we propose a modular agentic pipeline, MedRaC, that combines retrieval-augmented generation and Python-based code execution. Without any fine-tuning, MedRaC improves the accuracy of different LLMs from 16.35% up to 53.19%. Our work highlights the limitations of current benchmark practices and proposes a more clinically faithful methodology. By enabling transparent and transferable reasoning evaluation, we move closer to making LLM-based systems trustworthy for real-world medical applications.
Benlu Wang, Iris Xia, Junda Wang, Feiyun Ouyang, Arman Cohan, Hong Yu 0001, Zonghai Yao
EMNLP5
2025 MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
abstract
Automatic question generation (QG) is essential for AI and NLP, particularly in intelligent tutoring, dialogue systems, and fact verification. Generating multiple-choice questions (MCQG) for professional exams, like the United States Medical Licensing Examination (USMLE), is particularly challenging, requiring domain expertise and complex multi-hop reasoning for high-quality questions. However, current large language models (LLMs) like GPT-4 struggle with professional MCQG due to outdated knowledge, hallucination issues, and prompt sensitivity, resulting in unsatisfactory quality and difficulty. To address these challenges, we propose MCQG-SRefine, an LLM self-refine-based (Critique and Correction) framework for converting medical cases into high-quality USMLE-style questions. By integrating expert-driven prompt engineering with iterative self-critique and self-correction feedback, MCQG-SRefine significantly enhances human expert satisfaction regarding both the quality and difficulty of the questions. Furthermore, we introduce an LLM-as-Judge-based automatic metric to replace the complex and costly expert evaluation process, ensuring reliable and expert-aligned assessments.
Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang 0001, Hong Yu 0001
NAACL (Long Papers)5
2024 SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization
abstract
Large Language Models (LLMs) such as GPT & Llama have demonstrated significant achievements in summarization tasks but struggle with factual inaccuracies, a critical issue in clinical NLP applications where errors could lead to serious consequences.To counter the high costs and limited availability of expert-annotated data for factual alignment, this study introduces an innovative pipeline that utilizes >100B parameter GPT variants like GPT-3.5 & GPT-4 to act as synthetic experts to generate high-quality synthetics feedback aimed at enhancing factual consistency in clinical note summarization.Our research primarily focuses on edit feedback generated by these synthetic feedback experts without additional human annotations, mirroring and optimizing the practical scenario in which medical professionals refine AI system outputs.Although such 100B+ parameter GPT variants have proven to demonstrate expertise in various clinical NLP tasks, such as the Medical Licensing Examination, there is scant research on their capacity to act as synthetic feedback experts and deliver expert-level edit feedback for improving the generation quality of weaker (<10B parameter) LLMs like GPT-2 (1.5B) & Llama 2 (7B) in clinical domain.So in this work, we leverage 100B+ GPT variants to act as synthetic feedback experts offering expert-level edit feedback, that is used to reduce hallucinations and align weaker (<10B parameter) LLMs with medical facts using two distinct alignment algorithms (DPO & SALT), endeavoring to narrow the divide between AIgenerated content and factual accuracy.This highlights the substantial potential of LLMbased synthetic edits in enhancing the alignment of clinical factuality 1 . * indicates equal contribution † Presently in AMD AI
Prakamya Mishra, Zonghai Yao, Parth Vashisht, Feiyun Ouyang, Beining Wang, Vidhi Dhaval Mody, Hong Yu 0001
EMNLP4