Nishanth Anand

dblp:241/7250 · also Nishanth V. Anand · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
temporal difference learning
1.222023
Prediction and Control in Continual Reinforcement Learning · NeurIPS 2023
Preferential Temporal Difference Learning · ICML 2021
Machine learning › Reinforcement learning
value function estimation
1.222023
Prediction and Control in Continual Reinforcement Learning · NeurIPS 2023
Preferential Temporal Difference Learning · ICML 2021
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning
0.712023
Prediction and Control in Continual Reinforcement Learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 0.7complementary learning systems theory · 0.7state reweighting · 0.5linear function approximation · 0.5
YearPublicationVenuePosition
2023 Prediction and Control in Continual Reinforcement Learning
abstract
Temporal difference (TD) learning is often used to update the estimate of the value function which is used by RL agents to extract useful policies. In this paper, we focus on value function estimation in continual reinforcement learning. We propose to decompose the value function into two components which update at different timescales: a _permanent_ value function, which holds general knowledge that persists over time, and a _transient_ value function, which allows quick adaptation to new situations. We establish theoretical results showing that our approach is well suited for continual learning and draw connections to the complementary learning systems (CLS) theory from neuroscience. Empirically, this approach improves performance significantly on both prediction and control problems.
Nishanth Anand, Doina Precup
NeurIPS1
2021 Preferential Temporal Difference Learning
abstract
Temporal-Difference (TD) learning is a general and very useful tool for estimating the value function of a given policy, which in turn is required to find good policies. Generally speaking, TD learning updates states whenever they are visited. When the agent lands in a state, its value can be used to compute the TD-error, which is then propagated to other states. However, it may be interesting, when computing updates, to take into account other information than whether a state is visited or not. For example, some states might be more important than others (such as states which are frequently seen in a successful trajectory). Or, some states might have unreliable value estimates (for example, due to partial observability or lack of data), making their values less desirable as targets. We propose an approach to re-weighting states used in TD updates, both when they are the input and when they provide the target for the update. We prove that our approach converges with linear function approximation and illustrate its desirable empirical behaviour compared to other TD-style methods.
Nishanth Anand, Doina Precup
ICML1