EDBT 2026 Demo / reviewers in the wild / expert
Kihyun Yu
dblp:275/5513
· DBLP profile ↗
2ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.9 | 1 | 2025 | An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints · ICML 2025 |
Machine learning › Reinforcement learning › online decision making
online reinforcement learning |
0.9 | 1 | 2025 | An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints · ICML 2025 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.9 | 1 | 2025 | An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
primal-dual method · 0.9optimistic mirror descent · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Optimistic Algorithm for online CMDPS with Anytime Adversarial ConstraintsabstractOnline safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn optimal policies that maximize rewards while satisfying safety constraints modeled by constrained Markov decision processes (CMDPs). Existing methods achieve sublinear regret under stochastic constraints but often fail in adversarial settings, where constraints are unknown, time-varying, and potentially adversarially designed. In this paper, we propose the Optimistic Mirror Descent Primal-Dual (OMDPD) algorithm, the first to address online CMDPs with anytime adversarial constraints. OMDPD achieves optimal regret $\tilde{\mathcal{O}}(\sqrt{K})$ and strong constraint violation $\tilde{\mathcal{O}}(\sqrt{K})$ without relying on Slater’s condition or the existence of a strictly known safe policy. We further show that access to accurate estimates of rewards and transitions can further improve these bounds. Our results offer practical guarantees for safe decision-making in adversarial environments. Kihyun Yu, Dabeen Lee, Xin Liu 0049, Honghao Wei |
ICML | 2 |
| 2020 | Collaborative SLAM and AR-guided navigation for floor layout inspection
Kihyun Yu, Jeonghyeon Ahn |
Vis. Comput. | 1 |