VLDB 2026 Research / reviewers in the wild / expert
Arrasy Rahman
dblp:242/9328
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-1006-9653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 51% Multi-agent systems · 27% Efficient and distributed learning · 19% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
2.0 | 3 | 2025 | HyperMARL: Adaptive Hypernetworks for Multi-Agent RL · NeurIPS 2025 A General Learning Framework for Open Ad Hoc Teamwork Using Graph-based Policy Learning · J. Mach. Learn. Res. 2023 Scaling Multi-Agent Reinforcement Learning with Selective Parameter Sharing · ICML 2021 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
ad hoc teamwork |
1.9 | 3 | 2024 | N-agent Ad Hoc Teamwork · NeurIPS 2024 A General Learning Framework for Open Ad Hoc Teamwork Using Graph-based Policy Learning · J. Mach. Learn. Res. 2023 Towards Open Ad Hoc Teamwork Using Graph-based Policy Learning · ICML 2021 |
Machine learning › Efficient and distributed learning
parameter sharing |
1.4 | 2 | 2025 | HyperMARL: Adaptive Hypernetworks for Multi-Agent RL · NeurIPS 2025 Scaling Multi-Agent Reinforcement Learning with Selective Parameter Sharing · ICML 2021 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning |
0.8 | 1 | 2024 | N-agent Ad Hoc Teamwork · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent modeling |
0.5 | 1 | 2021 | Towards Open Ad Hoc Teamwork Using Graph-based Policy Learning · ICML 2021 |
Machine learning › Deep learning architectures and training
hypernetwork |
0.3 | 1 | 2025 | HyperMARL: Adaptive Hypernetworks for Multi-Agent RL · NeurIPS 2025 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.2 | 1 | 2024 | N-agent Ad Hoc Teamwork · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
scalable multi-agent reinforcement learning |
0.1 | 1 | 2021 | Scaling Multi-Agent Reinforcement Learning with Selective Parameter Sharing · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
policy gradient · 1.6graph neural network · 1.2hypernetwork · 0.9agent-conditioned parameter generation · 0.9teammate behavior representation learning · 0.8partial observability · 0.7belief estimation · 0.7parameter sharing · 0.5joint-action value learning · 0.5agent partitioning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HyperMARL: Adaptive Hypernetworks for Multi-Agent RLabstractAdaptive cooperation in multi-agent reinforcement learning (MARL) requires policies to express homogeneous, specialised, or mixed behaviours, yet achieving this adaptivity remains a critical challenge. While parameter sharing (PS) is standard for efficient learning, it notoriously suppresses the behavioural diversity required for specialisation. This failure is largely due to cross-agent gradient interference, a problem we find is surprisingly exacerbated by the common practice of *coupling agent IDs with observations*. Existing remedies typically add complexity through altered objectives, manual preset diversity levels, or sequential updates -- raising a fundamental question: *can shared policies adapt without these intricacies?* We propose a solution built on a key insight: an agent-conditioned hypernetwork can generate agent-specific parameters and *decouple* observation- and agent-conditioned gradients, directly countering the interference from coupling agent IDs with observations. Our resulting method, **HyperMARL**, avoids the complexities of prior work and empirically reduces policy gradient variance. Across diverse MARL benchmarks (22 scenarios, up to 30 agents), HyperMARL achieves performance competitive with six key baselines while preserving behavioural diversity comparable to non-parameter sharing methods, establishing it as a versatile and principled approach for adaptive MARL. The code is publicly available at https://github.com/KaleabTessera/HyperMARL. Kale-ab Abebe Tessera, Arrasy Rahman, Amos J. Storkey, Stefano V. Albrecht |
NeurIPS | 2 |
| 2024 | N-agent Ad Hoc TeamworkabstractCurrent approaches to learning cooperative multi-agent behaviors assume relatively restrictive settings. In standard fully cooperative multi-agent reinforcement learning, the learning algorithm controls *all* agents in the scenario, while in ad hoc teamwork, the learning algorithm usually assumes control over only a *single* agent in the scenario. However, many cooperative settings in the real world are much less restrictive. For example, in an autonomous driving scenario, a company might train its cars with the same learning algorithm, yet once on the road, these cars must cooperate with cars from another company. Towards expanding the class of scenarios that cooperative learning methods may optimally address, we introduce $N$*-agent ad hoc teamwork* (NAHT), where a set of autonomous agents must interact and cooperate with dynamically varying numbers and types of teammates. This paper formalizes the problem, and proposes the *Policy Optimization with Agent Modelling* (POAM) algorithm. POAM is a policy gradient, multi-agent reinforcement learning approach to the NAHT problem, that enables adaptation to diverse teammate behaviors by learning representations of teammate behaviors. Empirical evaluation on tasks from the multi-agent particle environment and StarCraft II shows that POAM improves cooperative task returns compared to baseline approaches, and enables out-of-distribution generalization to unseen teammates. Caroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman, Peter Stone 0001 |
NeurIPS | 2 |
| 2023 | A General Learning Framework for Open Ad Hoc Teamwork Using Graph-based Policy LearningabstractOpen ad hoc teamwork is the problem of training a single agent to efficiently collaborate with an unknown group of teammates whose composition may change over time. A variable team composition creates challenges for the agent, such as the requirement to adapt to new team dynamics and dealing with changing state vector sizes. These challenges are aggravated in real-world applications in which the controlled agent only has a partial view of the environment. In this work, we develop a class of solutions for open ad hoc teamwork under full and partial observability. We start by developing a solution for the fully observable case that leverages graph neural network architectures to obtain an optimal policy based on reinforcement learning. We then extend this solution to partially observable scenarios by proposing different methodologies that maintain belief estimates over the latent environment states and team composition. These belief estimates are combined with our solution for the fully observable case to compute an agent's optimal policy under partial observability in open ad hoc teamwork. Empirical results demonstrate that our solution can learn efficient policies in open ad hoc teamwork in fully and partially observable cases. Further analysis demonstrates that our methods' success is a result of effectively learning the effects of teammates' actions while also inferring the inherent state of the environment under partial observability. Arrasy Rahman, Ignacio Carlucho, Niklas Höpner, Stefano V. Albrecht |
J. Mach. Learn. Res. | 1 |
| 2022 | A Survey of Ad Hoc Teamwork Research
Reuth Mirsky, Ignacio Carlucho, Arrasy Rahman, Elliot Fosong, William Macke, Mohan Sridharan, Peter Stone 0001, Stefano V. Albrecht |
EUMAS | 3 |
| 2021 | Scaling Multi-Agent Reinforcement Learning with Selective Parameter SharingabstractSharing parameters in multi-agent deep reinforcement learning has played an essential role in allowing algorithms to scale to a large number of agents. Parameter sharing between agents significantly decreases the number of trainable parameters, shortening training times to tractable levels, and has been linked to more efficient learning. However, having all agents share the same parameters can also have a detrimental effect on learning. We demonstrate the impact of parameter sharing methods on training speed and converged returns, establishing that when applied indiscriminately, their effectiveness is highly dependent on the environment. We propose a novel method to automatically identify agents which may benefit from sharing parameters by partitioning them based on their abilities and goals. Our approach combines the increased sample efficiency of parameter sharing with the representational capacity of multiple independent networks to reduce training time and increase final returns. Filippos Christianos, Georgios Papoudakis, Arrasy Rahman, Stefano V. Albrecht |
ICML | 3 |
| 2021 | Towards Open Ad Hoc Teamwork Using Graph-based Policy LearningabstractAd hoc teamwork is the challenging problem of designing an autonomous agent which can adapt quickly to collaborate with teammates without prior coordination mechanisms, including joint training. Prior work in this area has focused on closed teams in which the number of agents is fixed. In this work, we consider open teams by allowing agents with different fixed policies to enter and leave the environment without prior notification. Our solution builds on graph neural networks to learn agent models and joint-action value models under varying team compositions. We contribute a novel action-value computation that integrates the agent model and joint-action value model to produce action-value estimates. We empirically demonstrate that our approach successfully models the effects other agents have on the learner, leading to policies that robustly adapt to dynamic team compositions and significantly outperform several alternative methods. Arrasy Rahman, Niklas Höpner, Filippos Christianos, Stefano V. Albrecht |
ICML | 1 |
| 2021 | Interpretable Goal Recognition in the Presence of Occluded Factors for Autonomous VehiclesabstractRecognising the goals or intentions of observed vehicles is a key step towards predicting the long-term future behaviour of other agents in an autonomous driving scenario. When there are unseen obstacles or occluded vehicles in a scenario, goal recognition may be confounded by the effects of these unseen entities on the behaviour of observed vehicles. Existing prediction algorithms that assume rational behaviour with respect to inferred goals may fail to make accurate long-horizon predictions because they ignore the possibility that the behaviour is influenced by such unseen entities. We introduce the Goal and Occluded Factor Inference (GOFI) algorithm which bases inference on inverse-planning to jointly infer a probabilistic belief over goals and potential occluded factors. We then show how these beliefs can be integrated into Monte Carlo Tree Search (MCTS). We demonstrate that jointly inferring goals and occluded factors leads to more accurate beliefs with respect to the true world state and allows an agent to safely navigate several scenarios where other baselines take unsafe actions leading to collisions. Josiah Hanna, Arrasy Rahman, Elliot Fosong, Francisco Girbal Eiras, Mihai Dobre, John Redford, Subramanian Ramamoorthy, Stefano V. Albrecht |
IROS | 2 |