EDBT 2026 Demo / reviewers in the wild / expert
Ruosen Li
dblp:351/0775
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Question answering and dialogue systems · 27% Language models and text generation · 23% Reinforcement learning · 18% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 6 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reward learning
reward modeling |
1.0 | 1 | 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
verification |
1.0 | 1 | 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
automatic evaluation |
0.8 | 1 | 2024 | IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems
interactive question answering |
0.8 | 1 | 2024 | IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.8 | 1 | 2024 | MEQA: A Benchmark for Multi-hop Event-centric Question Answering with Explanations · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model
reasoning model |
0.3 | 1 | 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
persona assignment · 1.5large language model · 1.5semi-automatic question generation · 0.8explanation evaluation metrics · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from ExperienceabstractZhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001 |
ACL (1) | 2 |
| 2024 | IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question AnsweringabstractTo evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on directly assessing the immediate responses generated by the models based on the given question and context. In the common use case of humans seeking AI assistant’s help in finding information, these non-interactive evaluations do not account for the dynamic nature of human-model conversations, and interaction-aware evaluations have shown that accurate models are not necessarily preferred by humans Lee et al. Recent works in human-computer interaction (HCI) have employed human evaluators to conduct interactions and evaluations, but they are often prohibitively expensive and time-consuming to scale. In this work, we introduce an automated evaluation framework IQA-EVAL to Interactive Question Answering Evaluations, more specifically, we introduce LLM-based Evaluation Agent (LEA) that can: (1) simulate human behaviors to generate interactions with IQA models; (2) automatically evaluate the generated interactions. Moreover, we propose assigning personas to LEAs to better simulate groups of real human evaluators. We show that: (1) our evaluation framework with GPT-4 (or Claude) as the backbone model achieves a high correlation with human evaluations on the IQA task; (2) assigning personas to LEA to better represent the crowd further significantly improves correlations. Finally, we use our automated metric to evaluate five recent LLMs with over 1000 questions from complex and ambiguous question answering tasks, which would cost $5k if evaluated by humans. Ruosen Li, Barry Wang, Xinya Du |
NeurIPS | 1 |
| 2024 | MEQA: A Benchmark for Multi-hop Event-centric Question Answering with ExplanationsabstractExisting benchmarks for multi-hop question answering (QA) primarily evaluate models based on their ability to reason about entities and the relationships between them. However, there's a lack of insight into how these models perform in terms of both events and entities. In this paper, we introduce a novel semi-automatic question generation strategy by composing event structures from information extraction (IE) datasets and present the first Multi-hop Event-centric Question Answering (MEQA) benchmark. It contains (1) 2,243 challenging questions that require a diverse range of complex reasoning over entity-entity, entity-event, and event-event relations; (2) corresponding multi-step QA-format event reasoning chain (explanation) which leads to the answer for each question. We also introduce two metrics for evaluating explanations: completeness and logical consistency. We conduct comprehensive benchmarking and analysis, which shows that MEQA is challenging for the latest state-of-the-art models encompassing large language models (LLMs); and how they fall short of providing faithful explanations of the event-centric reasoning process. Ruosen Li, Son Quoc Tran, Lei Xia 0004, Xinya Du |
NeurIPS | 1 |