EDBT 2026 Demo / reviewers in the wild / expert
Jia Lin Hau
dblp:329/5798
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
1.5 | 2 | 2025 | Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025 On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision Processes · NeurIPS 2023 |
Machine learning › Reinforcement learning
markov decision process |
0.9 | 2 | 2025 | On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision Processes · NeurIPS 2023 Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
entropic risk measure |
0.9 | 1 | 2025 | Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.9 | 1 | 2025 | Risk-Averse Total-Reward Reinforcement Learning · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
entropic value-at-risk · 1.5q-learning · 0.9dynamic consistency · 0.9value-at-risk · 0.7conditional value-at-risk · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence AnalysisabstractIn Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents’ preferences for certain outcomes. This paper proposes a new Q-learning algorithm for quantile optimization in MDPs with strong convergence and performance guarantees. The algorithm leverages a new, simple dynamic program (DP) decomposition for quantile MDPs. Compared with prior work, our DP decomposition requires neither known transition probabilities nor solving complex saddle point equations and serves as a suitable foundation for other model-free RL algorithms. Our numerical results in tabular domains show that our Q-learning algorithm converges to its DP variant and outperforms earlier algorithms. Jia Lin Hau, Erick Delage, Esther Derman, Mohammad Ghavamzadeh, Marek Petrik |
AISTATS | 1 |
| 2025 | Risk-Averse Total-Reward Reinforcement LearningabstractRisk-averse total-reward Markov Decision Processes (MDPs) offer a promising framework for modeling and solving undiscounted infinite-horizon objectives. Existing model-based algorithms for risk measures like the entropic risk measure (ERM) and entropic value-at-risk (EVaR) are effective in small problems, but require full access to transition probabilities. We propose a Q-learning algorithm to compute the optimal stationary policy for total-reward ERM and EVaR objectives with strong convergence and performance guarantees. The algorithm and its optimality are made possible by ERM's dynamic consistency and elicitability. Our numerical results on tabular domains demonstrate quick and reliable convergence of the proposed Q-learning algorithm to the optimal risk-averse value function. Xihong Su, Jia Lin Hau, Gersi Doko, Kishan Panaganti, Marek Petrik |
NeurIPS | 2 |
| 2023 | Entropic Risk Optimization in Discounted MDPsabstractRisk-averse Markov Decision Processes (MDPs) have optimal policies that achieve high returns with low variability, but these MDPs are often difficult to solve. Only a few practical risk-averse objectives admit a dynamic programming (DP) formulation, which is the mainstay of most MDP and RL algorithms. We derive a new DP formulation for discounted risk-averse MDPs with Entropic Risk Measure (ERM) and Entropic Value at Risk (EVaR) objectives. Our DP formulation for ERM, which is possible because of our novel definition of value function with time-dependent risk levels, can approximate optimal policies in a time that is polynomial in the approximation error. We then use the ERM algorithm to optimize the EVaR objective in polynomial time using an optimized discretization scheme. Our numerical results show the viability of our formulations and algorithms in discounted MDPs. Jia Lin Hau, Marek Petrik, Mohammad Ghavamzadeh |
AISTATS | 1 |
| 2023 | On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision ProcessesabstractOptimizing static risk-averse objectives in Markov decision processes is difficult because they do not admit standard dynamic programming equations common in Reinforcement Learning (RL) algorithms. Dynamic programming decompositions that augment the state space with discrete risk levels have recently gained popularity in the RL community. Prior work has shown that these decompositions are optimal when the risk level is discretized sufficiently. However, we show that these popular decompositions for Conditional-Value-at-Risk (CVaR) and Entropic-Value-at-Risk (EVaR) are inherently suboptimal regardless of the discretization level. In particular, we show that a saddle point property assumed to hold in prior literature may be violated. However, a decomposition does hold for Value-at-Risk and our proof demonstrates how this risk measure differs from CVaR and EVaR. Our findings are significant because risk-averse algorithms are used in high-stake environments, making their correctness much more critical. Jia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek Petrik |
NeurIPS | 1 |