VLDB 2026 Research / reviewers in the wild / expert
Manuel Kroiss
dblp:255/4852
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 67% Transfer learning and domain adaptation · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.4 | 1 | 2020 | What Can Learned Intrinsic Rewards Capture? · ICML 2020 |
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-gradient |
0.4 | 1 | 2020 | What Can Learned Intrinsic Rewards Capture? · ICML 2020 |
Machine learning › Reinforcement learning
reward learning |
0.4 | 1 | 2020 | What Can Learned Intrinsic Rewards Capture? · ICML 2020 |
Methods — techniques the papers use, named apart from their topics
meta-gradient reinforcement learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | What Can Learned Intrinsic Rewards Capture?abstractThe objective of a reinforcement learning agent is to behave so as to maximise the sum of a suitable scalar function of state: the reward. These rewards are typically given and immutable. In this paper, we instead consider the proposition that the reward function itself can be a good locus of learned knowledge. To investigate this, we propose a scalable meta-gradient framework for learning useful intrinsic reward functions across multiple lifetimes of experience. Through several proof-of-concept experiments, we show that it is feasible to learn and capture knowledge about long-term exploration and exploitation into a reward function. Furthermore, we show that unlike policy transfer methods that capture “how” the agent should behave, the learned reward functions can generalise to other kinds of agents and to changes in the dynamics of the environment by capturing “what” the agent should strive to do. Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
ICML | 5 |