EDBT 2026 Demo / reviewers in the wild / expert
Ruoyu Yao
dblp:288/9301
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-1754-1770ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 67% Multi-agent systems · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration |
0.9 | 1 | 2025 | HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
multimodal retrieval-augmented generation |
0.9 | 1 | 2025 | HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation · ACM Multimedia 2025 |
Information retrieval › distributed information retrieval
multi-source retrieval |
0.3 | 1 | 2025 | HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
query rewriting · 1.7large language model · 1.7consistency voting · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | nuPlan-R: A Closed-Loop Planning Benchmark for Autonomous Driving via Reactive Multi-Agent SimulationabstractRecent advances in closed-loop planning benchmarks have significantly improved the evaluation of autonomous vehicles. However, existing benchmarks still rely on rule-based reactive agents such as the Intelligent Driver Model (IDM), which lack behavioral diversity and fail to capture realistic human interactions, leading to oversimplified traffic dynamics. To address these limitations, we present nuPlan-R, a new reactive closed-loop planning benchmark that integrates learning-based reactive multi-agent simulation into the nuPlan framework. Our benchmark replaces the rule-based IDM agents with noise-decoupled diffusion-based reactive agents and introduces an interaction-aware agent selection mechanism to ensure both realism and computational efficiency. Furthermore, we extend the benchmark with two additional metrics to enable a more comprehensive assessment of planning performance. Extensive experiments demonstrate that our reactive agent model produces more realistic, diverse, and human-like traffic behaviors, leading to a benchmark environment that better reflects real-world interactive driving. We further reimplement a collection of rule-based, learning-based, and hybrid planning approaches within our nuPlan-R benchmark, providing a clearer reflection of planner performance in complex interactive scenarios and better highlighting the advantages of learning-based planners in handling complex and dynamic scenarios. These results establish nuPlan-R as a new standard for fair, reactive, and realistic closed-loop planning evaluation. We will open-source the code for the new benchmark. We have open-sourced our framework at https://github.com/Pemixing/nuPlan-R. Mingxing Peng, Ruoyu Yao, Xusen Guo, Jun Ma 0008 |
IV | 2 |
| 2025 | LMMCoDrive: Cooperative Driving with Large Multimodal ModelsabstractTo address the intricate challenges of cooperative scheduling and motion planning in Autonomous Mobility-on-Demand (AMoD) systems, this paper introduces LMMCoDrive, a novel cooperative driving framework that leverages a Large Multimodal Model (LMM) to improve traffic efficiency and passenger experience in dynamic urban environments. This framework seamlessly integrates scheduling and motion planning processes to ensure the effective operation of Cooperative Autonomous Vehicles (CAVs). The spatial relationship between CAVs and passenger requests is abstracted into a Bird’s-Eye View (BEV) image to fully exploit the potential of the multimodal understanding ability of LMMs. Besides, trajectories are cautiously refined for each CAV while ensuring collision avoidance through safety constraints. A decentralized optimization strategy, facilitated by the Alternating Direction Method of Multipliers (ADMM) within the LMM framework, is proposed to drive the graph evolution of CAVs. Simulation results in diverse urban scenarios demonstrate the pivotal role and significant impact of LMM in optimizing CAV scheduling and seamlessly serving a decentralized cooperative optimization process for each CAV. This marks a substantial stride towards practical, efficient, and safe AMoD systems that are poised to revolutionize urban transportation. The code is available at https://github.com/henryhcliu/LMMCoDrive. Haichao Liu 0003, Ruoyu Yao, Zhenmin Huang, Shaojie Shen, Jun Ma 0008 |
IROS | 2 |
| 2025 | HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented GenerationabstractWhile Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge, conventional single-agent RAG remains fundamentally limited in resolving complex queries demanding coordinated reasoning across heterogeneous data ecosystems. We present HM-RAG, a novel Hierarchical Multi-agent Multimodal RAG framework that pioneers collaborative intelligence for dynamic knowledge synthesis across structured, unstructured, and graph-based data. The framework is composed of a three-tiered architecture with specialized agents: a Decomposition Agent that dissects complex queries into contextually coherent sub-tasks via semantic-aware query rewriting and schema-guided context augmentation; Multi-source Retrieval Agents that carry out parallel, modality-specific retrieval using plug-and-play modules designed for vector, graph, and web-based databases; and a Decision Agent that uses consistency voting to integrate multi-source answers and resolve discrepancies in retrieval results through Expert Model Refinement. This architecture attains comprehensive query understanding by combining textual, graph-relational, and web-derived evidence, resulting in a remarkable 12.95% improvement in answer accuracy and a 3.56% boost in question classification accuracy over baseline RAG systems on the ScienceQA and CrisisMMD benchmarks. Notably, HM-RAG establishes state-of-the-art results in zero-shot settings on both datasets. Its modular architecture ensures seamless integration of new data modalities while maintaining strict data governance, marking a significant advancement in addressing the critical challenges of multimodal reasoning and knowledge synthesis in RAG systems. Ruoyu Yao, Siyuan Meng, Ding Wang 0001, Jun Ma 0008 |
ACM Multimedia | 3 |
| 2024 | Hierarchical Uncertainty-aware Autonomous Driving in Lane-changing Scenarios: Behavior Prediction and Motion PlanningabstractSafe and efficient interactions with surrounding vehicles in multilane driving are essential for autonomous vehicles. However, achieving smooth and flexible responses to surrounding vehicles’ lane changes remains a challenge due to the uncertainties in the behavior prediction progress. Deep learning-based methods were manifested powerful in modeling agents’ motion uncertainties for making stochastic intention classification and trajectory prediction. Nevertheless, performance degradation are likely to occur when the black-box model makes multi-modal predictions in unseen situations. This paper proposes a novel AV planning framework that combines deep learning-based behavior prediction and optimization-based uncertainty-aware motion planning to resolve these challenges. We hierarchically address uncertainties inherent in both behavior patterns and model performance through an adaptive motion planning approach, using an improved constrained iterative linear quadratic regulator that handles non-convex constraints and non-Gaussian uncertainties while minimizing travel costs. Evaluations using INTERACTION and HighD datasets demonstrate the effectiveness of uncertainty-aware planning in enhancing AV safety performance. Ruoyu Yao |
IV | 1 |