Jia Lin Hau

dblp:329/5798 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning
1.522025
Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025
On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision Processes · NeurIPS 2023
Machine learning › Reinforcement learning
markov decision process
0.922025
On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision Processes · NeurIPS 2023
Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
entropic risk measure
0.912025
Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.912025
Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

entropic value-at-risk · 1.5q-learning · 0.9dynamic consistency · 0.9value-at-risk · 0.7conditional value-at-risk · 0.7
YearPublicationVenuePosition
2025 Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
abstract
In Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents’ preferences for certain outcomes. This paper proposes a new Q-learning algorithm for quantile optimization in MDPs with strong convergence and performance guarantees. The algorithm leverages a new, simple dynamic program (DP) decomposition for quantile MDPs. Compared with prior work, our DP decomposition requires neither known transition probabilities nor solving complex saddle point equations and serves as a suitable foundation for other model-free RL algorithms. Our numerical results in tabular domains show that our Q-learning algorithm converges to its DP variant and outperforms earlier algorithms.
Jia Lin Hau, Erick Delage, Esther Derman, Mohammad Ghavamzadeh, Marek Petrik
AISTATS1
2025 Risk-Averse Total-Reward Reinforcement Learning
abstract
Risk-averse total-reward Markov Decision Processes (MDPs) offer a promising framework for modeling and solving undiscounted infinite-horizon objectives. Existing model-based algorithms for risk measures like the entropic risk measure (ERM) and entropic value-at-risk (EVaR) are effective in small problems, but require full access to transition probabilities. We propose a Q-learning algorithm to compute the optimal stationary policy for total-reward ERM and EVaR objectives with strong convergence and performance guarantees. The algorithm and its optimality are made possible by ERM's dynamic consistency and elicitability. Our numerical results on tabular domains demonstrate quick and reliable convergence of the proposed Q-learning algorithm to the optimal risk-averse value function.
Xihong Su, Jia Lin Hau, Gersi Doko, Kishan Panaganti, Marek Petrik
NeurIPS2
2023 Entropic Risk Optimization in Discounted MDPs
abstract
Risk-averse Markov Decision Processes (MDPs) have optimal policies that achieve high returns with low variability, but these MDPs are often difficult to solve. Only a few practical risk-averse objectives admit a dynamic programming (DP) formulation, which is the mainstay of most MDP and RL algorithms. We derive a new DP formulation for discounted risk-averse MDPs with Entropic Risk Measure (ERM) and Entropic Value at Risk (EVaR) objectives. Our DP formulation for ERM, which is possible because of our novel definition of value function with time-dependent risk levels, can approximate optimal policies in a time that is polynomial in the approximation error. We then use the ERM algorithm to optimize the EVaR objective in polynomial time using an optimized discretization scheme. Our numerical results show the viability of our formulations and algorithms in discounted MDPs.
Jia Lin Hau, Marek Petrik, Mohammad Ghavamzadeh
AISTATS1
2023 On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision Processes
abstract
Optimizing static risk-averse objectives in Markov decision processes is difficult because they do not admit standard dynamic programming equations common in Reinforcement Learning (RL) algorithms. Dynamic programming decompositions that augment the state space with discrete risk levels have recently gained popularity in the RL community. Prior work has shown that these decompositions are optimal when the risk level is discretized sufficiently. However, we show that these popular decompositions for Conditional-Value-at-Risk (CVaR) and Entropic-Value-at-Risk (EVaR) are inherently suboptimal regardless of the discretization level. In particular, we show that a saddle point property assumed to hold in prior literature may be violated. However, a decomposition does hold for Value-at-Risk and our proof demonstrates how this risk measure differs from CVaR and EVaR. Our findings are significant because risk-averse algorithms are used in high-stake environments, making their correctness much more critical.
Jia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek Petrik
NeurIPS1