EDBT 2026 Demo / reviewers in the wild / expert
Nishanth Anand
dblp:241/7250 · also Nishanth V. Anand
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
temporal difference learning |
1.2 | 2 | 2023 | Prediction and Control in Continual Reinforcement Learning · NeurIPS 2023 Preferential Temporal Difference Learning · ICML 2021 |
Machine learning › Reinforcement learning
value function estimation |
1.2 | 2 | 2023 | Prediction and Control in Continual Reinforcement Learning · NeurIPS 2023 Preferential Temporal Difference Learning · ICML 2021 |
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning |
0.7 | 1 | 2023 | Prediction and Control in Continual Reinforcement Learning · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
temporal difference learning · 0.7complementary learning systems theory · 0.7state reweighting · 0.5linear function approximation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Prediction and Control in Continual Reinforcement LearningabstractTemporal difference (TD) learning is often used to update the estimate of the value function which is used by RL agents to extract useful policies. In this paper, we focus on value function estimation in continual reinforcement learning. We propose to decompose the value function into two components which update at different timescales: a _permanent_ value function, which holds general knowledge that persists over time, and a _transient_ value function, which allows quick adaptation to new situations. We establish theoretical results showing that our approach is well suited for continual learning and draw connections to the complementary learning systems (CLS) theory from neuroscience. Empirically, this approach improves performance significantly on both prediction and control problems. Nishanth Anand, Doina Precup |
NeurIPS | 1 |
| 2021 | Preferential Temporal Difference LearningabstractTemporal-Difference (TD) learning is a general and very useful tool for estimating the value function of a given policy, which in turn is required to find good policies. Generally speaking, TD learning updates states whenever they are visited. When the agent lands in a state, its value can be used to compute the TD-error, which is then propagated to other states. However, it may be interesting, when computing updates, to take into account other information than whether a state is visited or not. For example, some states might be more important than others (such as states which are frequently seen in a successful trajectory). Or, some states might have unreliable value estimates (for example, due to partial observability or lack of data), making their values less desirable as targets. We propose an approach to re-weighting states used in TD updates, both when they are the input and when they provide the target for the update. We prove that our approach converges with linear function approximation and illustrate its desirable empirical behaviour compared to other TD-style methods. Nishanth Anand, Doina Precup |
ICML | 1 |