EDBT 2026 Demo / reviewers in the wild / expert
Chris Nota
dblp:236/4989
· DBLP profile ↗
3ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 75% Optimization for machine learning · 14% Learning paradigms · 12% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.9 | 2 | 2021 | Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods · ICML 2021 Asynchronous Coagent Networks · ICML 2020 |
Machine learning › Reinforcement learning
value function |
0.5 | 1 | 2021 | Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods · ICML 2021 |
Machine learning › Optimization for machine learning
variance reduction |
0.5 | 1 | 2021 | Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods · ICML 2021 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.4 | 1 | 2020 | Asynchronous Coagent Networks · ICML 2020 |
Machine learning › Learning paradigms
lifelong learning |
0.4 | 1 | 2020 | Lifelong Learning with a Changing Action Set · AAAI 2020 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option-critic |
0.4 | 1 | 2020 | Asynchronous Coagent Networks · ICML 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2020 | Lifelong Learning with a Changing Action Set · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
posterior inference · 0.5policy gradient · 0.5hindsight · 0.5structure inference · 0.4recurrent network · 0.4policy optimization · 0.4convergence analysis · 0.4asynchronous networks · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Posterior Value Functions: Hindsight Baselines for Policy Gradient MethodsabstractHindsight allows reinforcement learning agents to leverage new observations to make inferences about earlier states and transitions. In this paper, we exploit the idea of hindsight and introduce posterior value functions. Posterior value functions are computed by inferring the posterior distribution over hidden components of the state in previous timesteps and can be used to construct novel unbiased baselines for policy gradient methods. Importantly, we prove that these baselines reduce (and never increase) the variance of policy gradient estimators compared to traditional state value functions. While the posterior value function is motivated by partial observability, we extend these results to arbitrary stochastic MDPs by showing that hindsight-capable agents can model stochasticity in the environment as a special case of partial observability. Finally, we introduce a pair of methods for learning posterior value functions and prove their convergence. Chris Nota, Philip S. Thomas, Bruno C. da Silva 0001 |
ICML | 1 |
| 2020 | Lifelong Learning with a Changing Action SetabstractIn many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transition dynamics, changing rewards functions, etc. have been well-studied in the lifelong learning literature, the setting where the size of the action set changes remains unaddressed. In this paper, we present first steps towards developing an algorithm that autonomously adapts to an action set whose size changes over time. To tackle this open problem, we break it into two problems that can be solved iteratively: inferring the underlying, unknown, structure in the space of actions and optimizing a policy that leverages this structure. We demonstrate the efficiency of this approach on large-scale real-world lifelong learning problems. Yash Chandak, Georgios Theocharous, Chris Nota, Philip S. Thomas |
AAAI | 3 |
| 2020 | Asynchronous Coagent NetworksabstractCoagent policy gradient algorithms (CPGAs) are reinforcement learning algorithms for training a class of stochastic neural networks called coagent networks. In this work, we prove that CPGAs converge to locally optimal policies. Additionally, we extend prior theory to encompass asynchronous and recurrent coagent networks. These extensions facilitate the straightforward design and analysis of hierarchical reinforcement learning algorithms like the option-critic, and eliminate the need for complex derivations of customized learning rules for these algorithms. James E. Kostas, Chris Nota, Philip S. Thomas |
ICML | 2 |