Anuj Mahajan

dblp:99/3800 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2229-5287ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 72% Planning, search and constraint satisfaction · 11% Probabilistic and Bayesian machine learning · 6%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 20 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.552023
SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning · NeurIPS 2023
Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021
UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
1.532021
Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021
UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021
RODE: Learning Roles to Decompose Multi-Agent Tasks · ICLR 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
1.022023
SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning · NeurIPS 2023
MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019
Machine learning › Reinforcement learning
exploration
0.922021
UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021
MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019
Knowledge, reasoning and agents › Multi-agent systems
decentralized planning
0.812024
Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
language-model-based planning
0.812024
Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent planning
0.812024
Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments · NeurIPS 2024
Machine learning › Reinforcement learning
partial observability
0.712023
SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.622021
Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021
MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019
Machine learning › Reinforcement learning
bellman equation
0.512021
Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › exploration › multi-robot exploration
coordinated exploration
0.512021
UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning
actor-critic methods
0.412019
VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.412019
VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning
policy optimization
0.412019
VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning › exploration › exploration strategies
temporally-extended exploration
0.412019
MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.412019
VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › statistical relational learning
lifted inference
0.212015
Lifted Inference Rules With Constraints · NIPS 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning
statistical relational learning
0.212015
Lifted Inference Rules With Constraints · NIPS 2015
Information theory › estimation theory › bayesian estimation
MAP inference
0.212015
Lifted Inference Rules With Constraints · NIPS 2015
Machine learning › Learning theory › PAC learning
PAC analysis
0.112021
Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021

Methods — techniques the papers use, named apart from their topics

plan-act-correct-verify · 0.8large language model · 0.8centralised training with decentralised execution · 0.7value decomposition · 0.5universal successor features · 0.5tensor decomposition · 0.5low-rank tensor approximation · 0.5PAC analysis · 0.5variational inference · 0.4KL divergence · 0.4single occurrence rule · 0.2generalized binomial rule · 0.2decomposer rule · 0.2constraint language · 0.2
YearPublicationVenuePosition
2025 UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents
Harsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan, Ari Seff, Di Feng, Ruijia Cheng, Andres Romero Mier Y. Teran, Esteban Gomez, Abhishek Sundararajan, Forrest Huang, Amanda Swearngin, Mohana Prasad Sathya Moorthy, Jeffrey Nichols 0001, Alexander Toshev
ICCV4
2025 From Interaction to Impact: Towards Safer AI Agent Through Understanding and Evaluating Mobile UI Operation Impacts
Zhuohao (Jerry) Zhang, Eldon Schoop, Jeffrey Nichols 0001, Anuj Mahajan, Amanda Swearngin
IUI4
2024 Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments
abstract
The ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditional planning methods that rely on domain-specific knowledge and handcrafted rules, LMs generalize from diverse data and adapt to various tasks with minimal tuning, acting as a compressed knowledge base. However, LMs in their standard form face challenges with long-horizon tasks, particularly in partially observable multi-agent settings. We propose an LM-based Long-Horizon Planner for Multi-Agent Robotics (LLaMAR), a cognitive architecture for planning that achieves state-of-the-art results in long-horizon tasks within partially observable environments. LLaMAR employs a plan-act-correct-verify framework, allowing self-correction from action execution feedback without relying on oracles or simulators. Additionally, we present MAP-THOR, a comprehensive test suite encompassing household tasks of varying complexity within the AI2-THOR environment. Experiments show that LLaMAR achieves a 30\% higher success rate than other state-of-the-art LM-based multi-agent planners in MAP-THOR and Search \& Rescue tasks. Code can be found at [https://github.com/nsidn98/LLaMAR](https://github.com/nsidn98/LLaMAR)
Siddharth Nayak, Adelmo Morrison Orozco, Marina Ten Have, Jackson Zhang, Vittal Thirumalai, Darren Chen, Aditya Kapoor, Eric Robinson, Karthik Gopalakrishnan 0002, James Harrison, Anuj Mahajan, Brian Ichter, Hamsa Balakrishnan
NeurIPS11
2023 Effects of Spectral Normalization in Multi-Agent Reinforcement Learning
abstract
A reliable critic is central to on-policy actor-critic learning. But it becomes challenging to learn a reliable critic in a multi-agent sparse reward scenario due to two factors: 1) The joint action space grows exponentially with the number of agents 2) This, combined with the reward sparseness and environment noise, leads to large sample requirements for accurate learning. We show that regularising the critic with spectral normalization (SN) enables it to learn more robustly, even in multi-agent on-policy sparse reward scenarios. Our experiments show that the regularised critic is quickly able to learn from the sparse rewarding experience in the complex SMAC and RWARE domains. These findings highlight the importance of regularisation in the critic for stable learning.
Kinal Mehta, Anuj Mahajan, Pawan Kumar 0001
IJCNN2
2023 SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning
abstract
The availability of challenging benchmarks has played a key role in the recent progress of machine learning. In cooperative multi-agent reinforcement learning, the StarCraft Multi-Agent Challenge (SMAC) has become a popular testbed for centralised training with decentralised execution. However, after years of sustained improvement on SMAC, algorithms now achieve near-perfect performance. In this work, we conduct new analysis demonstrating that SMAC lacks the stochasticity and partial observability to require complex closed-loop policies. In particular, we show that an open-loop policy conditioned only on the timestep can achieve non-trivial win rates for many SMAC scenarios. To address this limitation, we introduce SMACv2, a new version of the benchmark where scenarios are procedurally generated and require agents to generalise to previously unseen settings (from the same distribution) during evaluation. We also introduce the extended partial observability challenge (EPO), which augments SMACv2 to ensure meaningful partial observability. We show that these changes ensure the benchmarkrequires the use of closed-loop policies. We evaluate state-of-the-art algorithms on SMACv2 and show that it presents significant challenges not present in the original benchmark. Our analysis illustrates that SMACv2 addresses the discovered deficiencies of SMAC and can help benchmark the next generation of MARL methods. Videos of training are available on our website.
Benjamin Ellis, Jonathan Cook 0004, Skander Moalla, Mikayel Samvelyan, Mingfei Sun 0001, Anuj Mahajan, Jakob N. Foerster, Shimon Whiteson
NeurIPS6
2021 RODE: Learning Roles to Decompose Multi-Agent Tasks
Tonghan Wang 0001, Tarun Gupta 0002, Anuj Mahajan, Bei Peng 0001, Shimon Whiteson, Chongjie Zhang
ICLR3
2021 UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning
abstract
VDN and QMIX are two popular value-based algorithms for cooperative MARL that learn a centralized action value function as a monotonic mixing of per-agent utilities. While this enables easy decentralization of the learned policy, the restricted joint action value function can prevent them from solving tasks that require significant coordination between agents at a given timestep. We show that this problem can be overcome by improving the joint exploration of all agents during training. Specifically, we propose a novel MARL approach called Universal Value Exploration (UneVEn) that learns a set of related tasks simultaneously with a linear decomposition of universal successor features. With the policies of already solved related tasks, the joint exploration process of all agents can be improved to help them achieve better coordination. Empirical results on a set of exploration games, challenging cooperative predator-prey tasks requiring significant coordination among agents, and StarCraft II micromanagement benchmarks show that UneVEn can solve tasks where other state-of-the-art MARL methods fail.
Tarun Gupta 0002, Anuj Mahajan, Bei Peng 0001, Wendelin Böhmer, Shimon Whiteson
ICML2
2021 Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning
abstract
Reinforcement Learning in large action spaces is a challenging problem. This is especially true for cooperative multi-agent reinforcement learning (MARL), which often requires tractable learning while respecting various constraints like communication budget and information about other agents. In this work, we focus on the fundamental hurdle affecting both value-based and policy-gradient approaches: an exponential blowup of the action space with the number of agents. For value-based methods, it poses challenges in accurately representing the optimal value function for value-based methods, thus inducing suboptimality. For policy gradient methods, it renders the critic ineffective and exacerbates the problem of the lagging critic. We show that from a learning theory perspective, both problems can be addressed by accurately representing the associated action-value function with a low-complexity hypothesis class. This requires accurately modelling the agent interactions in a sample efficient way. To this end, we propose a novel tensorised formulation of the Bellman equation. This gives rise to our method Tesseract, which utilises the view of Q-function seen as a tensor where the modes correspond to action spaces of different agents. Algorithms derived from Tesseract decompose the Q-tensor across the agents and utilise low-rank tensor approximations to model the agent interactions relevant to the task. We provide PAC analysis for Tesseract based algorithms and highlight their relevance to the class of rich observation MDPs. Empirical results in different domains confirm the gains in sample efficiency using Tesseract as supported by the theory.
Anuj Mahajan, Mikayel Samvelyan, Viktor Makoviychuk, Animesh Garg, Jean Kossaifi, Shimon Whiteson, Yuke Zhu, Anima Anandkumar
ICML1
2019 VIREL: A Variational Inference Framework for Reinforcement Learning
abstract
Applying probabilistic models to reinforcement learning (RL) enables the uses of powerful optimisation tools such as variational inference in RL. However, existing inference frameworks and their algorithms pose significant challenges for learning optimal policies, e.g., the lack of mode capturing behaviour in pseudo-likelihood methods, difficulties learning deterministic policies in maximum entropy RL based approaches, and a lack of analysis when function approximators are used. We propose VIREL, a theoretically grounded probabilistic inference framework for RL that utilises a parametrised action-value function to summarise future dynamics of the underlying MDP, generalising existing approaches. VIREL also benefits from a mode-seeking form of KL divergence, the ability to learn deterministic optimal polices naturally from inference, and the ability to optimise value functions and policies in separate, iterative steps. In applying variational expectation-maximisation to VIREL, we thus show that the actor-critic algorithm can be reduced to expectation-maximisation, with policy improvement equivalent to an E-step and policy evaluation to an M-step. We then derive a family of actor-critic methods fromVIREL, including a scheme for adaptive exploration. Finally, we demonstrate that actor-critic algorithms from this family outperform state-of-the-art methods based on soft value functions in several domains.
Matthew Fellows, Anuj Mahajan, Tim G. J. Rudner, Shimon Whiteson
NeurIPS2
2019 MAVEN: Multi-Agent Variational Exploration
abstract
Centralised training with decentralised execution is an important setting for cooperative deep multi-agent reinforcement learning due to communication constraints during execution and computational tractability in training. In this paper, we analyse value-based methods that are known to have superior performance in complex environments. We specifically focus on QMIX, the current state-of-the-art in this domain. We show that the representation constraints on the joint action-values introduced by QMIX and similar methods lead to provably poor exploration and suboptimality. Furthermore, we propose a novel approach called MAVEN that hybridises value and policy-based methods by introducing a latent space for hierarchical control. The value-based agents condition their behaviour on the shared latent variable controlled by a hierarchical policy. This allows MAVEN to achieve committed, temporally extended exploration, which is key to solving complex multi-agent tasks. Our experimental results show that MAVEN achieves significant performance improvements on the challenging SMAC domain.
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, Shimon Whiteson
NeurIPS1
2015 Feature Selection for Short Text Classification using Wavelet Packet Transform
abstract
Text classification tasks suffer from curse of dimensionality due to large feature space.Short text data further exacerbates the problem due to their sparse and noisy nature.Feature selection thus becomes an important step in improving the classification performance.In this paper, we propose a novel feature selection method using Wavelet Packet Transform.Wavelet Packet Transform (WPT) has been used widely in various fields due to its efficiency in encoding transient signals.We demonstrate how short text classification task can be benefited by feature selection using WPT due to their sparse nature.Our technique chooses the most discriminating features by computing inter-class distances in the transformed space.We experimented extensively with several short text datasets.Compared to well known techniques our approach reduces the feature space size and improves the overall classification performance significantly in all the datasets.
Anuj Mahajan, Sharmistha Jat, Shourya Roy
CoNLL1
2015 Lifted Inference Rules With Constraints
abstract
Lifted inference rules exploit symmetries for fast reasoning in statistical rela-tional models. Computational complexity of these rules is highly dependent onthe choice of the constraint language they operate on and therefore coming upwith the right kind of representation is critical to the success of lifted inference.In this paper, we propose a new constraint language, called setineq, which allowssubset, equality and inequality constraints, to represent substitutions over the vari-ables in the theory. Our constraint formulation is strictly more expressive thanexisting representations, yet easy to operate on. We reformulate the three mainlifting rules: decomposer, generalized binomial and the recently proposed singleoccurrence for MAP inference, to work with our constraint representation. Exper-iments on benchmark MLNs for exact and sampling based inference demonstratethe effectiveness of our approach over several other existing techniques.
Happy Mittal, Anuj Mahajan, Vibhav Gogate, Parag Singla
NIPS2
2008 Mining Financial News for Major Events and Their Impacts on the Market
abstract
In this paper we have proposed a stock market analysis system that analyzes financial news items to identify and characterize major events that impact the market. The events have been identified using latent Dirichlet allocation (LDA) based topic extraction mechanism. These topics have been thereafter analyzed in conjunction with actual market data to understand their impact on the market. A prediction system has been proposed which can predict whether the stock market will fall or rise, based on news items.
Anuj Mahajan, Lipika Dey, S. K. Mirajul Haque
Web Intelligence1