Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yiheng Bing

dblp:421/0504 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 87% Probabilistic and Bayesian machine learning · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
constrained reinforcement learning
0.912025
Extreme Value Policy Optimization for Safe Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning
safe reinforcement learning
0.912025
Extreme Value Policy Optimization for Safe Reinforcement Learning · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning
extreme value theory
0.312025
Extreme Value Policy Optimization for Safe Reinforcement Learning · ICML 2025

Methods — techniques the papers use, named apart from their topics

extreme quantile optimization · 0.9extreme prioritization replay · 0.9
YearPublicationVenuePosition
2025 Extreme Value Policy Optimization for Safe Reinforcement Learning
abstract
Ensuring safety is a critical challenge in applying Reinforcement Learning (RL) to real-world scenarios. Constrained Reinforcement Learning (CRL) addresses this by maximizing returns under predefined constraints, typically formulated as the expected cumulative cost. However, expectation-based constraints overlook rare but high-impact extreme value events in the tail distribution, such as black swan incidents, which can lead to severe constraint violations. To address this issue, we propose the Extreme Value policy Optimization (EVO) algorithm, leveraging Extreme Value Theory (EVT) to model and exploit extreme reward and cost samples, reducing constraint violations. EVO introduces an extreme quantile optimization objective to explicitly capture extreme samples in the cost tail distribution. Additionally, we propose an extreme prioritization mechanism during replay, amplifying the learning signal from rare but high-impact extreme samples. Theoretically, we establish upper bounds on expected constraint violations during policy updates, guaranteeing strict constraint satisfaction at a zero-violation quantile level. Further, we demonstrate that EVO achieves a lower probability of constraint violations than expectation-based methods and exhibits lower variance than quantile regression methods. Extensive experiments show that EVO significantly reduces constraint violations during training while maintaining competitive policy performance compared to baselines.
Shiqing Gao, Yihang Zhou, Haoyu Luo, Yiheng Bing, Jiaxin Ding 0001, Luoyi Fu, Xinbing Wang
ICML5