VLDB 2026 Research / reviewers in the wild / expert
Yuqi Bian
dblp:343/9107
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 92% Generative modeling · 8% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
1.7 | 2 | 2025 | Learning to Reuse Policies in State Evolvable Environments · ICML 2025 Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory Stitching · ICLR 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory Stitching · ICLR 2025 |
Machine learning › Reinforcement learning
policy adaptation |
0.9 | 1 | 2025 | Learning to Reuse Policies in State Evolvable Environments · ICML 2025 |
Machine learning › Reinforcement learning › ensemble reinforcement learning
policy ensemble |
0.9 | 1 | 2025 | Learning to Reuse Policies in State Evolvable Environments · ICML 2025 |
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy reuse |
0.9 | 1 | 2025 | Learning to Reuse Policies in State Evolvable Environments · ICML 2025 |
Machine learning › Reinforcement learning › offline reinforcement learning
trajectory stitching |
0.9 | 1 | 2025 | Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory Stitching · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation |
0.3 | 1 | 2025 | Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory Stitching · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory Stitching · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
state reconstruction · 0.9offline learning · 0.9diffusion model · 0.9credit assignment · 0.9bidirectional dynamics constraint · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory StitchingabstractLearning from offline data without interacting with the environment is a promising way to fully leverage the intelligent decision-making capabilities of multi-agent reinforcement learning (MARL). Previous approaches have primarily focused on developing learning techniques, such as conservative methods tailored to MARL using limited offline data. However, these methods often overlook the temporal relationships across different timesteps and spatial relationships between teammates, resulting in low learning efficiency in imbalanced data scenarios. To comprehensively explore the data structure of MARL and enhance learning efficiency, we propose Multi-Agent offline coordination via Diffusion-based Trajectory Stitching (MADiTS), a novel diffusion-based data augmentation pipeline that systematically generates trajectories by stitching high-quality coordination segments together. MADiTS first generates trajectory segments using a trained diffusion model, followed by applying a bidirectional dynamics constraint to ensure that the trajectories align with environmental dynamics. Additionally, we develop an offline credit assignment technique to identify and optimize the behavior of underperforming agents in the generated segments. This iterative procedure continues until a satisfactory augmented episode trajectory is generated within the predefined limit or is discarded otherwise. Empirical results on imbalanced datasets of multiple benchmarks demonstrate that MADiTS significantly improves MARL performance. Lei Yuan 0005, Yuqi Bian, Lihe Li, Cong Guan, Yang Yu 0001 |
ICLR | 2 |
| 2025 | Learning to Reuse Policies in State Evolvable EnvironmentsabstractThe policy trained via reinforcement learning (RL) makes decisions based on sensor-derived state features. It is common for state features to evolve for reasons such as periodic sensor maintenance or the addition of new sensors for performance improvement. The deployed policy fails in new state space when state features are unseen during training. Previous work tackles this challenge by training a sensor-invariant policy or generating multiple policies and selecting the appropriate one with limited samples. However, both directions struggle to guarantee the performance when faced with unpredictable evolutions. In this paper, we formalize this problem as state evolvable reinforcement learning (SERL), where the agent is required to mitigate policy degradation after state evolutions without costly exploration. We propose Lapse by reusing policies learned from the old state space in two distinct aspects. On one hand, Lapse directly reuses the robust old policy by composing it with a learned state reconstruction model to handle vanishing sensors. On the other hand, the behavioral experience from the old policy is reused by Lapse to train a newly adaptive policy through offline learning, better utilizing new sensors. To leverage advantages of both policies in different scenarios, we further propose automatic ensemble weight adjustment to effectively aggregate them. Theoretically, we justify that robust policy reuse helps mitigate uncertainty and error from both evolution and reconstruction. Empirically, Lapse achieves a significant performance improvement, outperforming the strongest baseline by about $2\times$ in benchmark environments. Bohan Yang 0018, Lihe Li, Yuqi Bian, Ruiqi Xue, Feng Chen 0042, Yi-Chen Li 0001, Lei Yuan 0005, Yang Yu 0001 |
ICML | 4 |