Chris Nota

dblp:236/4989 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 75% Optimization for machine learning · 14% Learning paradigms · 12%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.922021
Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods · ICML 2021
Asynchronous Coagent Networks · ICML 2020
Machine learning › Reinforcement learning
value function
0.512021
Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods · ICML 2021
Machine learning › Optimization for machine learning
variance reduction
0.512021
Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods · ICML 2021
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412020
Asynchronous Coagent Networks · ICML 2020
Machine learning › Learning paradigms
lifelong learning
0.412020
Lifelong Learning with a Changing Action Set · AAAI 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option-critic
0.412020
Asynchronous Coagent Networks · ICML 2020
Machine learning › Reinforcement learning
policy optimization
0.412020
Lifelong Learning with a Changing Action Set · AAAI 2020

Methods — techniques the papers use, named apart from their topics

posterior inference · 0.5policy gradient · 0.5hindsight · 0.5structure inference · 0.4recurrent network · 0.4policy optimization · 0.4convergence analysis · 0.4asynchronous networks · 0.4
YearPublicationVenuePosition
2021 Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods
abstract
Hindsight allows reinforcement learning agents to leverage new observations to make inferences about earlier states and transitions. In this paper, we exploit the idea of hindsight and introduce posterior value functions. Posterior value functions are computed by inferring the posterior distribution over hidden components of the state in previous timesteps and can be used to construct novel unbiased baselines for policy gradient methods. Importantly, we prove that these baselines reduce (and never increase) the variance of policy gradient estimators compared to traditional state value functions. While the posterior value function is motivated by partial observability, we extend these results to arbitrary stochastic MDPs by showing that hindsight-capable agents can model stochasticity in the environment as a special case of partial observability. Finally, we introduce a pair of methods for learning posterior value functions and prove their convergence.
Chris Nota, Philip S. Thomas, Bruno C. da Silva 0001
ICML1
2020 Lifelong Learning with a Changing Action Set
abstract
In many real-world sequential decision making problems, the number of available actions (decisions) can vary over time. While problems like catastrophic forgetting, changing transition dynamics, changing rewards functions, etc. have been well-studied in the lifelong learning literature, the setting where the size of the action set changes remains unaddressed. In this paper, we present first steps towards developing an algorithm that autonomously adapts to an action set whose size changes over time. To tackle this open problem, we break it into two problems that can be solved iteratively: inferring the underlying, unknown, structure in the space of actions and optimizing a policy that leverages this structure. We demonstrate the efficiency of this approach on large-scale real-world lifelong learning problems.
Yash Chandak, Georgios Theocharous, Chris Nota, Philip S. Thomas
AAAI3
2020 Asynchronous Coagent Networks
abstract
Coagent policy gradient algorithms (CPGAs) are reinforcement learning algorithms for training a class of stochastic neural networks called coagent networks. In this work, we prove that CPGAs converge to locally optimal policies. Additionally, we extend prior theory to encompass asynchronous and recurrent coagent networks. These extensions facilitate the straightforward design and analysis of hierarchical reinforcement learning algorithms like the option-critic, and eliminate the need for complex derivations of customized learning rules for these algorithms.
James E. Kostas, Chris Nota, Philip S. Thomas
ICML2