EDBT 2026 Demo / reviewers in the wild / expert
Mahdi Alikhasi
dblp:382/3605
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.8 | 1 | 2024 | Unveiling Options with Neural Network Decomposition · ICLR 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.8 | 1 | 2024 | Unveiling Options with Neural Network Decomposition · ICLR 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction |
0.8 | 1 | 2024 | Unveiling Options with Neural Network Decomposition · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
neural network decomposition · 0.8levin loss minimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unveiling Options with Neural Network DecompositionabstractIn reinforcement learning, agents often learn policies for specific tasks without the ability to generalize this knowledge to related tasks. This paper introduces an algorithm that attempts to address this limitation by decomposing neural networks encoding policies for Markov Decision Processes into reusable sub-policies, which are used to synthesize temporally extended actions, or options. We consider neural networks with piecewise linear activation functions, so that they can be mapped to an equivalent tree that is similar to oblique decision trees. Since each node in such a tree serves as a function of the input of the tree, each sub-tree is a sub-policy of the main policy. We turn each of these sub-policies into options by wrapping it with while-loops of varied number of iterations. Given the large number of options, we propose a selection mechanism based on minimizing the Levin loss for a uniform policy on these options. Empirical results in two grid-world domains where exploration can be difficult confirm that our method can identify useful options, thereby accelerating the learning process on similar but different tasks. Mahdi Alikhasi, Levi Lelis |
ICLR | 1 |