Kangrui Ruan

dblp:324/0593 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0000-7850-0206ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 69% Multi-agent systems · 11% Question answering and dialogue systems · 11%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › imitation learning
causal imitation learning
2.032024
Causal Imitation for Markov Decision Processes: a Partial Identification Approach · NeurIPS 2024
Causal Imitation Learning via Inverse Reinforcement Learning · ICLR 2023
Learning Human Driving Behaviors with Sequential Causal Imitation Learning · AAAI 2022
Machine learning › Reinforcement learning
imitation learning
2.032024
Causal Imitation for Markov Decision Processes: a Partial Identification Approach · NeurIPS 2024
Causal Imitation Learning via Inverse Reinforcement Learning · ICLR 2023
Learning Human Driving Behaviors with Sequential Causal Imitation Learning · AAAI 2022
Knowledge, reasoning and agents › Multi-agent systems › agentic AI
agentic reasoning
1.012026
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
1.012026
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization · ACL (1) 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization · ACL (1) 2026
Machine learning › Reinforcement learning › imitation learning › generative imitation learning
generative adversarial imitation learning
0.812024
Causal Imitation for Markov Decision Processes: a Partial Identification Approach · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect identification
partial identification
0.812024
Causal Imitation for Markov Decision Processes: a Partial Identification Approach · NeurIPS 2024
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.712023
Causal Imitation Learning via Inverse Reinforcement Learning · ICLR 2023
Robotics › Autonomous driving
driver behavior modeling
0.212022
Learning Human Driving Behaviors with Sequential Causal Imitation Learning · AAAI 2022

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.0group policy optimization · 1.0partial identification · 0.8GAIL · 0.8causal inference · 0.7causal graphical criterion · 0.6backdoor adjustment · 0.6adversarial imitation learning · 0.6
YearPublicationVenuePosition
2026 Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
abstract
Yifeng Ding, Hung Le, Songyang Han, Kangrui Ruan, Zhenghui Jin, Varun Kumar, Zijian Wang, Anoop Deoras. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Songyang Han, Kangrui Ruan, Zhenghui Jin, Zijian Wang 0002, Anoop Deoras
ACL (1)4
2025 R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
abstract
Artificial intelligence has made significant strides in medical visual question answering (Med-VQA), yet prevalent studies often interpret images holistically, overlooking the visual regions of interest that may contain crucial information, potentially aligning with a doctor’s prior knowledge that can be incorporated with minimal annotations (e.g., bounding boxes). To address this gap, this paper introduces R-LLaVA, designed to enhance biomedical VQA understanding by integrating simple medical annotations as prior knowledge directly into the image space through CLIP. These annotated visual regions of interest are then fed into the LLaVA model during training, aiming to enrich the model’s understanding of biomedical queries. Experimental evaluation on four standard Med-VQA datasets demonstrates R-LLaVA’s superiority over existing state-of-the-art (SoTA) methods. Additionally, to verify the model’s capability in visual comprehension, a novel multiple-choice medical visual understanding dataset is introduced, confirming the positive impact of focusing on visual regions of interest in advancing biomedical VQA understanding.
Xupeng Chen, Zhixin Lai, Kangrui Ruan, Shichu Chen, Zuozhu Liu
IJCNN3
2024 S2E: Towards an End-to-End Entity Resolution Solution from Acoustic Signal
abstract
Traditional cascading Entity Resolution (ER) pipeline suffers from propagated errors from upstream tasks. We address this issue by formulating a new end-to-end (E2E) ER problem, Signal-to-Entity (S2E), resolving query entity mentions to actionable entities in textual catalogs directly from audio queries instead of audio transcriptions in raw or parsed format. Additionally, we extend the E2E Spoken Language Understanding framework by introducing a novel dimension to ER research. We adapt three public datasets for the S2E task, and propose a novel solution, which aligns the multimodal signals via an effective retrieval co-attention mechanism and refined multimodal objectives. Despite 42% smaller in terms of the total model size, the proposed design outperforms the cascading baseline by 2.6%, 47.0%, and 73.3% across the three datasets respectively with different acoustic conditions.
Kangrui Ruan, Jiyang Wang, Helian Feng, Ali Kebarighotbi
ICASSP1
2024 Causal Imitation for Markov Decision Processes: a Partial Identification Approach
abstract
Imitation learning enables an agent to learn from expert demonstrations when the performance measure is unknown and the reward signal is not specified. Standard imitation methods do not generally apply when the learner and the expert's sensory capabilities mismatch and demonstrations are contaminated with unobserved confounding bias. To address these challenges, recent advancements in causal imitation learning have been pursued. However, these methods often require access to underlying causal structures that might not always be available, posing practical challenges. In this paper, we investigate robust imitation learning within the framework of canonical Markov Decision Processes (MDPs) using partial identification, allowing the agent to achieve expert performance even when the system dynamics are not uniquely determined from the confounded expert demonstrations. Specifically, first, we theoretically demonstrate that when unobserved confounders (UCs) exist in an MDP, the learner is generally unable to imitate expert performance. We then explore imitation learning in partially identifiable settings --- either transition distribution or reward function is non-identifiable from the available data and knowledge. Augmenting the celebrated GAIL method (Ho \& Ermon, 2016), our analysis leads to two novel causal imitation algorithms that can obtain effective policies guaranteed to achieve expert performance.
Kangrui Ruan, Junzhe Zhang 0001, Xuan Di, Elias Bareinboim
NeurIPS1
2023 Causal Imitation Learning via Inverse Reinforcement Learning
Kangrui Ruan, Junzhe Zhang 0001, Xuan Di, Elias Bareinboim
ICLR1
2022 Learning Human Driving Behaviors with Sequential Causal Imitation Learning
abstract
Learning human driving behaviors is an efficient approach for self-driving vehicles. Traditional Imitation Learning (IL) methods assume that the expert demonstrations follow Markov Decision Processes (MDPs). However, in reality, this assumption does not always hold true. Spurious correlation may exist through the paths of historical variables because of the existence of unobserved confounders. Accounting for the latent causal relationships from unobserved variables to outcomes, this paper proposes Sequential Causal Imitation Learning (SeqCIL) for imitating driver behaviors. We develop a sequential causal template that generalizes the default MDP settings to one with Unobserved Confounders (MDPUC-HD). Then we develop a sufficient graphical criterion to determine when ignoring causality leads to poor performances in MDPUC-HD. Through the framework of Adversarial Imitation Learning, we develop a procedure to imitate the expert policy by blocking π-backdoor paths at each time step. Our methods are evaluated on a synthetic dataset and a real-world highway driving dataset, both demonstrating that the proposed procedure significantly outperforms non-causal imitation learning methods.
Kangrui Ruan, Xuan Di
AAAI1