Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ruosen Li

dblp:351/0775 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Question answering and dialogue systems · 27% Language models and text generation · 23% Reinforcement learning · 18%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reward learning
reward modeling
1.012026
Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience · ACL (1) 2026
Machine learning › Trustworthy machine learning
verification
1.012026
Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model evaluation
automatic evaluation
0.812024
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering · NeurIPS 2024
Natural language and speech › Question answering and dialogue systems
interactive question answering
0.812024
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering · NeurIPS 2024
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.812024
MEQA: A Benchmark for Multi-hop Event-centric Question Answering with Explanations · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model
reasoning model
0.312026
Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

persona assignment · 1.5large language model · 1.5semi-automatic question generation · 0.8explanation evaluation metrics · 0.8
YearPublicationVenuePosition
2026 Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
abstract
Zhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001
ACL (1)2
2024 IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
abstract
To evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on directly assessing the immediate responses generated by the models based on the given question and context. In the common use case of humans seeking AI assistant’s help in finding information, these non-interactive evaluations do not account for the dynamic nature of human-model conversations, and interaction-aware evaluations have shown that accurate models are not necessarily preferred by humans Lee et al. Recent works in human-computer interaction (HCI) have employed human evaluators to conduct interactions and evaluations, but they are often prohibitively expensive and time-consuming to scale. In this work, we introduce an automated evaluation framework IQA-EVAL to Interactive Question Answering Evaluations, more specifically, we introduce LLM-based Evaluation Agent (LEA) that can: (1) simulate human behaviors to generate interactions with IQA models; (2) automatically evaluate the generated interactions. Moreover, we propose assigning personas to LEAs to better simulate groups of real human evaluators. We show that: (1) our evaluation framework with GPT-4 (or Claude) as the backbone model achieves a high correlation with human evaluations on the IQA task; (2) assigning personas to LEA to better represent the crowd further significantly improves correlations. Finally, we use our automated metric to evaluate five recent LLMs with over 1000 questions from complex and ambiguous question answering tasks, which would cost $5k if evaluated by humans.
Ruosen Li, Barry Wang, Xinya Du
NeurIPS1
2024 MEQA: A Benchmark for Multi-hop Event-centric Question Answering with Explanations
abstract
Existing benchmarks for multi-hop question answering (QA) primarily evaluate models based on their ability to reason about entities and the relationships between them. However, there's a lack of insight into how these models perform in terms of both events and entities. In this paper, we introduce a novel semi-automatic question generation strategy by composing event structures from information extraction (IE) datasets and present the first Multi-hop Event-centric Question Answering (MEQA) benchmark. It contains (1) 2,243 challenging questions that require a diverse range of complex reasoning over entity-entity, entity-event, and event-event relations; (2) corresponding multi-step QA-format event reasoning chain (explanation) which leads to the answer for each question. We also introduce two metrics for evaluating explanations: completeness and logical consistency. We conduct comprehensive benchmarking and analysis, which shows that MEQA is challenging for the latest state-of-the-art models encompassing large language models (LLMs); and how they fall short of providing faithful explanations of the event-centric reasoning process.
Ruosen Li, Son Quoc Tran, Lei Xia 0004, Xinya Du
NeurIPS1