VLDB 2026 Research / reviewers in the wild / expert
Feiyun Ouyang
dblp:359/6097
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-7061-7351ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 70% Multi-agent systems · 7% Knowledge representation and reasoning · 7% | |
| Human-computer interaction and pervasive computing
1 paper |
Health and well-being technologies · 50% Human-AI interaction · 50% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
2.7 | 3 | 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026 From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations · EMNLP 2025 RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
knowledge-grounded reasoning |
1.0 | 1 | 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model reasoning |
1.0 | 1 | 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent reasoning |
1.0 | 1 | 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
multi-step planning |
1.0 | 1 | 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning · AAAI 2026 |
Health and well-being technologies › behavior change
health behavior change |
1.0 | 1 | 2026 | ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care · AAAI 2026 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality |
0.9 | 1 | 2025 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering |
0.9 | 1 | 2025 | From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
reasoning enhancement |
0.9 | 1 | 2025 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation › text summarization › biomedical summarization
clinical summarization |
0.8 | 1 | 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
factual consistency |
0.8 | 1 | 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation
text summarization |
0.8 | 1 | 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization · EMNLP 2024 |
Medical and health informatics › digital health
diabetes management |
0.3 | 1 | 2026 | ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care · AAAI 2026 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations · EMNLP 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.6persuasive strategy · 2.0multi-agent simulation · 2.0retrieval · 1.0multi-agent framework · 1.0chain-of-thought · 1.0sub-question decomposition · 0.9retrieval-augmented generation · 0.9retrieval augmentation · 0.9factuality scoring · 0.9code execution · 0.9DPO · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRIME: Planning and Retrieval-Integrated Memory for Enhanced ReasoningabstractInspired by the dual-process theory of human cognition from Thinking, Fast and Slow, we introduce PRIME (Planning and Retrieval-Integrated Memory for Enhanced Reasoning), a multi-agent reasoning framework that dynamically integrates System 1 (fast, intuitive thinking) and System 2 (slow, deliberate thinking). PRIME first employs a Quick Thinking Agent to generate a rapid answer; if uncertainty is detected, it then triggers a structured System 2 reasoning pipeline composed of specialized agents for planning, hypothesis generation, retrieval, information integration, and decision-making. This multi-agent design mimics human cognitive processes faithfully and enhances both efficiency and accuracy. Experimental results with LLaMA 3 models demonstrate that PRIME enables open-source LLMs to perform competitively with state-of-the-art closed-source models like GPT-4 and GPT-4o on benchmarks requiring multi-hop and knowledge-grounded reasoning. This research establishes PRIME as a scalable solution for improving LLMs in domains requiring complex, knowledge-intensive reasoning. Hieu Tran, Zonghai Yao, Nguyen Luong Tran, Zhichao Yang 0001, Feiyun Ouyang, Razieh Rahimi, Hong Yu 0001 |
AAAI | 5 |
| 2026 | ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes CareabstractReal-world adoption of closed-loop insulin delivery systems (CLIDS) in type 1 diabetes remains low, driven not by technical failure, but by diverse behavioral, psychosocial, and social barriers. We introduce ChatCLIDS, the first benchmark to rigorously evaluate LLM–driven persuasive dialogue for health behavior change. Our framework features a library of expert-validated virtual patients, each with clinically grounded, heterogeneous profiles and realistic adoption barriers, and simulates multi-turn interactions with nurse agents equipped with a diverse set of evidence-based persuasive strategies. ChatCLIDS uniquely supports longitudinal counseling and adversarial social influence scenarios, enabling robust, multi-dimensional evaluation. Our findings reveal that while larger and more reflective LLMs adapt strategies over time, all models struggle to overcome resistance, especially under realistic social pressure. These results highlight critical limitations of current LLMs for behavior change, and offer a high-fidelity, scalable testbed for advancing trustworthy persuasive AI in healthcare and beyond. Zonghai Yao, Talha Chafekar, Junda Wang, Feiyun Ouyang, Junhui Qian, Hong Yu 0001 |
AAAI | 5 |
| 2025 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Modelsabstract, which leverages information retrieval specifically for generated sub-questions and re-answers these sub-questions with the relevant contextual information. Additionally, a Retrieval-Augmented Factuality Scorer is proposed to replace the original discriminator, prioritizing reasoning paths that meet high standards of factuality. Experimental results with LLaMA 3.1 show that RARE enables open-source LLMs to achieve competitive performance with top closed-source models like GPT-4 and GPT-4o. This research establishes RARE as a scalable solution for improving LLMs in domains where logical coherence and factual integrity are critical. Hieu Tran, Zonghai Yao, Zhichao Yang 0001, Junda Wang, Feiyun Ouyang, Hong Yu 0001 |
ACL (1) | 7 |
| 2025 | From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical CalculationsabstractLarge language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical decision-making, remains underexplored and poorly evaluated. Existing benchmarks often assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. In this work, we revisit medical calculation evaluation with a stronger focus on clinical trustworthiness. First, we clean and restructure the MedCalc-Bench dataset and propose a new step-by-step evaluation pipeline that independently assesses formula selection, entity extraction, and arithmetic computation. Under this granular framework, the accuracy of GPT-4o drops from 62.7% to 43.6%, revealing errors masked by prior evaluations. Second, we introduce an automatic error analysis framework that generates structured attribution for each failure mode. Human evaluation confirms its alignment with expert judgment, enabling scalable and explainable diagnostics. Finally, we propose a modular agentic pipeline, MedRaC, that combines retrieval-augmented generation and Python-based code execution. Without any fine-tuning, MedRaC improves the accuracy of different LLMs from 16.35% up to 53.19%. Our work highlights the limitations of current benchmark practices and proposes a more clinically faithful methodology. By enabling transparent and transferable reasoning evaluation, we move closer to making LLM-based systems trustworthy for real-world medical applications. Benlu Wang, Iris Xia, Junda Wang, Feiyun Ouyang, Arman Cohan, Hong Yu 0001, Zonghai Yao |
EMNLP | 5 |
| 2025 | MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison FeedbackabstractAutomatic question generation (QG) is essential for AI and NLP, particularly in intelligent tutoring, dialogue systems, and fact verification. Generating multiple-choice questions (MCQG) for professional exams, like the United States Medical Licensing Examination (USMLE), is particularly challenging, requiring domain expertise and complex multi-hop reasoning for high-quality questions. However, current large language models (LLMs) like GPT-4 struggle with professional MCQG due to outdated knowledge, hallucination issues, and prompt sensitivity, resulting in unsatisfactory quality and difficulty. To address these challenges, we propose MCQG-SRefine, an LLM self-refine-based (Critique and Correction) framework for converting medical cases into high-quality USMLE-style questions. By integrating expert-driven prompt engineering with iterative self-critique and self-correction feedback, MCQG-SRefine significantly enhances human expert satisfaction regarding both the quality and difficulty of the questions. Furthermore, we introduce an LLM-as-Judge-based automatic metric to replace the complex and costly expert evaluation process, ensuring reliable and expert-aligned assessments. Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang 0001, Hong Yu 0001 |
NAACL (Long Papers) | 5 |
| 2024 | SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical SummarizationabstractLarge Language Models (LLMs) such as GPT & Llama have demonstrated significant achievements in summarization tasks but struggle with factual inaccuracies, a critical issue in clinical NLP applications where errors could lead to serious consequences.To counter the high costs and limited availability of expert-annotated data for factual alignment, this study introduces an innovative pipeline that utilizes >100B parameter GPT variants like GPT-3.5 & GPT-4 to act as synthetic experts to generate high-quality synthetics feedback aimed at enhancing factual consistency in clinical note summarization.Our research primarily focuses on edit feedback generated by these synthetic feedback experts without additional human annotations, mirroring and optimizing the practical scenario in which medical professionals refine AI system outputs.Although such 100B+ parameter GPT variants have proven to demonstrate expertise in various clinical NLP tasks, such as the Medical Licensing Examination, there is scant research on their capacity to act as synthetic feedback experts and deliver expert-level edit feedback for improving the generation quality of weaker (<10B parameter) LLMs like GPT-2 (1.5B) & Llama 2 (7B) in clinical domain.So in this work, we leverage 100B+ GPT variants to act as synthetic feedback experts offering expert-level edit feedback, that is used to reduce hallucinations and align weaker (<10B parameter) LLMs with medical facts using two distinct alignment algorithms (DPO & SALT), endeavoring to narrow the divide between AIgenerated content and factual accuracy.This highlights the substantial potential of LLMbased synthetic edits in enhancing the alignment of clinical factuality 1 . * indicates equal contribution † Presently in AMD AI Prakamya Mishra, Zonghai Yao, Parth Vashisht, Feiyun Ouyang, Beining Wang, Vidhi Dhaval Mody, Hong Yu 0001 |
EMNLP | 4 |