Hadi Nekoei

dblp:270/0435 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-9755-6659ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 64% Language models and text generation · 17% Efficient and distributed learning · 9%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.422025
A Generalist Hanabi Agent · ICLR 2025
Continuous Coordination As a Realistic Scenario for Lifelong Learning · ICML 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
0.912025
A Generalist Hanabi Agent · ICLR 2025
Machine learning › Reinforcement learning
generalist agents
0.912025
A Generalist Hanabi Agent · ICLR 2025
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Efficient and distributed learning › distillation
teacher-student distillation
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Natural language and speech › Language models and text generation › LLM agents › web agents
web agent training
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Learning paradigms
lifelong learning
0.512021
Continuous Coordination As a Realistic Scenario for Lifelong Learning · ICML 2021
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
zero-shot coordination
0.512021
Continuous Coordination As a Realistic Scenario for Lifelong Learning · ICML 2021
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412020
The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning · NeurIPS 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning
non-stationarity
0.112021
Continuous Coordination As a Realistic Scenario for Lifelong Learning · ICML 2021
Machine learning › Reinforcement learning
sample efficiency
0.112020
The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

recurrent replay · 0.9hyperparameter sensitivity analysis · 0.9distributed DQN · 0.9bootstrapping · 0.9multi-agent reinforcement learning · 0.5deep reinforcement learning · 0.5local change adaptation regret · 0.4
YearPublicationVenuePosition
2025 A Generalist Hanabi Agent
abstract
Traditional multi-agent reinforcement learning (MARL) systems can develop cooperative strategies through repeated interactions. However, these systems are unable to perform well on any other setting than the one they have been trained on, and struggle to successfully cooperate with unfamiliar collaborators. This is particularly visible in the Hanabi benchmark, a popular 2-to-5 player cooperative card-game which requires complex reasoning and precise assistance to other agents. Current MARL agents for Hanabi can only learn one specific game-setting (e.g., 2-player games), and play with the same algorithmic agents. This is in stark contrast to humans, who can quickly adjust their strategies to work with unfamiliar partners or situations. In this paper, we introduce Recurrent Replay Relevance Distributed DQN (R3D2), a generalist agent for Hanabi, designed to overcome these limitations. We reformulate the task using text, as language has been shown to improve transfer. We then propose a distributed MARL algorithm that copes with the resulting dynamic observation- and action-space. In doing so, our agent is the first that can play all game settings concurrently, and extend strategies learned from one setting to other ones. As a consequence, our agent also demonstrates the ability to collaborate with different algorithmic agents ---agents that are themselves unable to do so.
Arjun Vaithilingam Sudhakar, Hadi Nekoei, Mathieu Reymond, Janarthanan Rajendran, Sarath Chandar
ICLR2
2025 How to Train Your LLM Web Agent: A Statistical Diagnosis
abstract
Large language model (LLM) agents for web interfaces have advanced rapidly, yet open-source systems still lag behind proprietary agents. Bridging this gap is key to enabling customizable, efficient, and privacy-preserving agents. Two challenges hinder progress: the reproducibility issues in RL and LLM agent training, where results often depend on sensitive factors like seeds and decoding parameters, and the focus of prior work on single-step tasks, overlooking the complexities of web-based, multi-step decision-making. We address these gaps by providing a statistically driven study of training LLM agents for web tasks. Our two-stage pipeline combines imitation learning from a Llama 3.3 70B teacher with on-policy fine-tuning via Group Relative Policy Optimization (GRPO) on a Llama 3.1 8B student. Through 240 configuration sweeps and rigorous bootstrapping, we chart the first compute allocation curve for open-source LLM web agents. Our findings show that dedicating one-third of compute to teacher traces and the rest to RL improves MiniWoB++ success by 6 points and closes 60\% of the gap to GPT-4o on WorkArena, while cutting GPU costs by 45\%. We introduce a principled hyperparameter sensitivity analysis, offering actionable guidelines for robust and cost-effective agent training.
Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Thibault Le Sellier de Chezelles, Megh Thakkar, Nicolas Angelard-Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Steve (Xue) Liu, Alexandre Drouin, Alexandre Piché, Alexandre Lacoste, Massimo Caccia
NeurIPS4
2024 Correction to: Multi-agent reinforcement learning for fast-timescale demand response of residential loads
Vincent Mai, Philippe Maisonneuve, Hadi Nekoei, Liam Paull, Antoine Lesage-Landry
Mach. Learn.4
2024 Multi-agent reinforcement learning for fast-timescale demand response of residential loads
Vincent Mai, Philippe Maisonneuve, Hadi Nekoei, Liam Paull, Antoine Lesage-Landry
Mach. Learn.4
2021 Continuous Coordination As a Realistic Scenario for Lifelong Learning
abstract
Current deep reinforcement learning (RL) algorithms are still highly task-specific and lack the ability to generalize to new environments. Lifelong learning (LLL), however, aims at solving multiple tasks sequentially by efficiently transferring and using knowledge between tasks. Despite a surge of interest in lifelong RL in recent years, the lack of a realistic testbed makes robust evaluation of LLL algorithms difficult. Multi-agent RL (MARL), on the other hand, can be seen as a natural scenario for lifelong RL due to its inherent non-stationarity, since the agents’ policies change over time. In this work, we introduce a multi-agent lifelong learning testbed that supports both zero-shot and few-shot settings. Our setup is based on Hanabi {—} a partially-observable, fully cooperative multi-agent game that has been shown to be challenging for zero-shot coordination. Its large strategy space makes it a desirable environment for lifelong RL tasks. We evaluate several recent MARL methods, and benchmark state-of-the-art LLL algorithms in limited memory and computation regimes to shed light on their strengths and weaknesses. This continual learning paradigm also provides us with a pragmatic way of going beyond centralized training which is the most commonly used training protocol in MARL. We empirically show that the agents trained in our setup are able to coordinate well with unseen agents, without any additional assumptions made by previous works. The code and all pre-trained models are available at https://github.com/chandar-lab/Lifelong-Hanabi.
Hadi Nekoei, Akilesh Badrinaaraayanan, Aaron C. Courville, Sarath Chandar
ICML1
2020 The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning
abstract
Deep model-based Reinforcement Learning (RL) has the potential to substantially improve the sample-efficiency of deep RL. While various challenges have long held it back, a number of papers have recently come out reporting success with deep model-based methods. This is a great development, but the lack of a consistent metric to evaluate such methods makes it difficult to compare various approaches. For example, the common single-task sample-efficiency metric conflates improvements due to model-based learning with various other aspects, such as representation learning, making it difficult to assess true progress on model-based RL. To address this, we introduce an experimental setup to evaluate model-based behavior of RL methods, inspired by work from neuroscience on detecting model-based behavior in humans and animals. Our metric based on this setup, the Local Change Adaptation (LoCA) regret, measures how quickly an RL method adapts to a local change in the environment. Our metric can identify model-based behavior, even if the method uses a poor representation and provides insight in how close a method's behavior is from optimal model-based behavior. We use our setup to evaluate the model-based behavior of MuZero on a variation of the classic Mountain Car task.
Harm van Seijen, Hadi Nekoei, Evan Racah, Sarath Chandar
NeurIPS2