EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyu Wang 0018
dblp:58/4775-18
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-1587-6307ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 63% Planning, search and constraint satisfaction · 25% Probabilistic and Bayesian machine learning · 12% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
1.5 | 2 | 2025 | Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens · ICML 2025 Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization · ICLR 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
bayesian planning |
0.9 | 1 | 2025 | Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens · ICML 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
model-based planning |
0.9 | 1 | 2025 | Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens · ICML 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty |
0.9 | 1 | 2025 | Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference |
0.9 | 1 | 2025 | Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens · ICML 2025 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization · ICLR 2023 |
Machine learning › Reinforcement learning
policy optimization |
0.7 | 1 | 2023 | Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization · ICLR 2023 |
Machine learning › Reinforcement learning › model-based reinforcement learning
value expansion |
0.7 | 1 | 2023 | Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
marginalization · 0.9doubly bayesian inference · 0.9belief updating · 0.9conservative bayesian model-based value expansion · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian LensabstractOffline reinforcement learning (RL) is crucial when online exploration is costly or unsafe but often struggles with high epistemic uncertainty due to limited data. Existing methods rely on fixed conservative policies, restricting adaptivity and generalization. To address this, we propose Reflect-then-Plan (RefPlan), a novel _doubly Bayesian_ offline model-based (MB) planning approach. RefPlan unifies uncertainty modeling and MB planning by recasting planning as Bayesian posterior estimation. At deployment, it updates a belief over environment dynamics using real-time observations, incorporating uncertainty into MB planning via marginalization. Empirical results on standard benchmarks show that RefPlan significantly improves the performance of conservative offline RL policies. In particular, RefPlan maintains robust performance under high epistemic uncertainty and limited data, while demonstrating resilience to changing environment dynamics, improving the flexibility, generalizability, and robustness of offline-learned policies. Jihwan Jeong, Xiaoyu Wang 0018, Jingmin Wang, Scott Sanner, Pascal Poupart |
ICML | 2 |
| 2024 | eMARLIN+: Addressing Partial Observability to Promote Traffic Signal Coordination by Leveraging Historical Information
Xiaoyu Wang 0018, Ayal Taitler, Ilia Smirnov, Scott Sanner, Baher Abdulhai |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization
Jihwan Jeong, Xiaoyu Wang 0018, Michael Gimelfarb, Baher Abdulhai, Scott Sanner |
ICLR | 2 |