VLDB 2026 Research / reviewers in the wild / expert
Anuj Mahajan
dblp:99/3800
· DBLP profile ↗
13ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2229-5287ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 72% Planning, search and constraint satisfaction · 11% Probabilistic and Bayesian machine learning · 6% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 20 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
2.5 | 5 | 2023 | SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning · NeurIPS 2023 Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021 UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
1.5 | 3 | 2021 | Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021 UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021 RODE: Learning Roles to Decompose Multi-Agent Tasks · ICLR 2021 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning |
1.0 | 2 | 2023 | SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning · NeurIPS 2023 MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019 |
Machine learning › Reinforcement learning
exploration |
0.9 | 2 | 2021 | UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021 MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019 |
Knowledge, reasoning and agents › Multi-agent systems
decentralized planning |
0.8 | 1 | 2024 | Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
language-model-based planning |
0.8 | 1 | 2024 | Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent planning |
0.8 | 1 | 2024 | Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments · NeurIPS 2024 |
Machine learning › Reinforcement learning
partial observability |
0.7 | 1 | 2023 | SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.6 | 2 | 2021 | Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021 MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019 |
Machine learning › Reinforcement learning
bellman equation |
0.5 | 1 | 2021 | Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › exploration › multi-robot exploration
coordinated exploration |
0.5 | 1 | 2021 | UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
actor-critic methods |
0.4 | 1 | 2019 | VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization |
0.4 | 1 | 2019 | VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2019 | VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning › exploration › exploration strategies
temporally-extended exploration |
0.4 | 1 | 2019 | MAVEN: Multi-Agent Variational Exploration · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.4 | 1 | 2019 | VIREL: A Variational Inference Framework for Reinforcement Learning · NeurIPS 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › statistical relational learning
lifted inference |
0.2 | 1 | 2015 | Lifted Inference Rules With Constraints · NIPS 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
statistical relational learning |
0.2 | 1 | 2015 | Lifted Inference Rules With Constraints · NIPS 2015 |
Information theory › estimation theory › bayesian estimation
MAP inference |
0.2 | 1 | 2015 | Lifted Inference Rules With Constraints · NIPS 2015 |
Machine learning › Learning theory › PAC learning
PAC analysis |
0.1 | 1 | 2021 | Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
plan-act-correct-verify · 0.8large language model · 0.8centralised training with decentralised execution · 0.7value decomposition · 0.5universal successor features · 0.5tensor decomposition · 0.5low-rank tensor approximation · 0.5PAC analysis · 0.5variational inference · 0.4KL divergence · 0.4single occurrence rule · 0.2generalized binomial rule · 0.2decomposer rule · 0.2constraint language · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents
Harsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan, Ari Seff, Di Feng, Ruijia Cheng, Andres Romero Mier Y. Teran, Esteban Gomez, Abhishek Sundararajan, Forrest Huang, Amanda Swearngin, Mohana Prasad Sathya Moorthy, Jeffrey Nichols 0001, Alexander Toshev |
ICCV | 4 |
| 2025 | From Interaction to Impact: Towards Safer AI Agent Through Understanding and Evaluating Mobile UI Operation Impacts
Zhuohao (Jerry) Zhang, Eldon Schoop, Jeffrey Nichols 0001, Anuj Mahajan, Amanda Swearngin |
IUI | 4 |
| 2024 | Long-Horizon Planning for Multi-Agent Robots in Partially Observable EnvironmentsabstractThe ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditional planning methods that rely on domain-specific knowledge and handcrafted rules, LMs generalize from diverse data and adapt to various tasks with minimal tuning, acting as a compressed knowledge base. However, LMs in their standard form face challenges with long-horizon tasks, particularly in partially observable multi-agent settings. We propose an LM-based Long-Horizon Planner for Multi-Agent Robotics (LLaMAR), a cognitive architecture for planning that achieves state-of-the-art results in long-horizon tasks within partially observable environments. LLaMAR employs a plan-act-correct-verify framework, allowing self-correction from action execution feedback without relying on oracles or simulators. Additionally, we present MAP-THOR, a comprehensive test suite encompassing household tasks of varying complexity within the AI2-THOR environment. Experiments show that LLaMAR achieves a 30\% higher success rate than other state-of-the-art LM-based multi-agent planners in MAP-THOR and Search \& Rescue tasks. Code can be found at [https://github.com/nsidn98/LLaMAR](https://github.com/nsidn98/LLaMAR) Siddharth Nayak, Adelmo Morrison Orozco, Marina Ten Have, Jackson Zhang, Vittal Thirumalai, Darren Chen, Aditya Kapoor, Eric Robinson, Karthik Gopalakrishnan 0002, James Harrison, Anuj Mahajan, Brian Ichter, Hamsa Balakrishnan |
NeurIPS | 11 |
| 2023 | Effects of Spectral Normalization in Multi-Agent Reinforcement LearningabstractA reliable critic is central to on-policy actor-critic learning. But it becomes challenging to learn a reliable critic in a multi-agent sparse reward scenario due to two factors: 1) The joint action space grows exponentially with the number of agents 2) This, combined with the reward sparseness and environment noise, leads to large sample requirements for accurate learning. We show that regularising the critic with spectral normalization (SN) enables it to learn more robustly, even in multi-agent on-policy sparse reward scenarios. Our experiments show that the regularised critic is quickly able to learn from the sparse rewarding experience in the complex SMAC and RWARE domains. These findings highlight the importance of regularisation in the critic for stable learning. Kinal Mehta, Anuj Mahajan, Pawan Kumar 0001 |
IJCNN | 2 |
| 2023 | SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement LearningabstractThe availability of challenging benchmarks has played a key role in the recent progress of machine learning. In cooperative multi-agent reinforcement learning, the StarCraft Multi-Agent Challenge (SMAC) has become a popular testbed for centralised training with decentralised execution. However, after years of sustained improvement on SMAC, algorithms now achieve near-perfect performance. In this work, we conduct new analysis demonstrating that SMAC lacks the stochasticity and partial observability to require complex closed-loop policies. In particular, we show that an open-loop policy conditioned only on the timestep can achieve non-trivial win rates for many SMAC scenarios. To address this limitation, we introduce SMACv2, a new version of the benchmark where scenarios are procedurally generated and require agents to generalise to previously unseen settings (from the same distribution) during evaluation. We also introduce the extended partial observability challenge (EPO), which augments SMACv2 to ensure meaningful partial observability. We show that these changes ensure the benchmarkrequires the use of closed-loop policies. We evaluate state-of-the-art algorithms on SMACv2 and show that it presents significant challenges not present in the original benchmark. Our analysis illustrates that SMACv2 addresses the discovered deficiencies of SMAC and can help benchmark the next generation of MARL methods. Videos of training are available on our website. Benjamin Ellis, Jonathan Cook 0004, Skander Moalla, Mikayel Samvelyan, Mingfei Sun 0001, Anuj Mahajan, Jakob N. Foerster, Shimon Whiteson |
NeurIPS | 6 |
| 2021 | RODE: Learning Roles to Decompose Multi-Agent Tasks
Tonghan Wang 0001, Tarun Gupta 0002, Anuj Mahajan, Bei Peng 0001, Shimon Whiteson, Chongjie Zhang |
ICLR | 3 |
| 2021 | UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningabstractVDN and QMIX are two popular value-based algorithms for cooperative MARL that learn a centralized action value function as a monotonic mixing of per-agent utilities. While this enables easy decentralization of the learned policy, the restricted joint action value function can prevent them from solving tasks that require significant coordination between agents at a given timestep. We show that this problem can be overcome by improving the joint exploration of all agents during training. Specifically, we propose a novel MARL approach called Universal Value Exploration (UneVEn) that learns a set of related tasks simultaneously with a linear decomposition of universal successor features. With the policies of already solved related tasks, the joint exploration process of all agents can be improved to help them achieve better coordination. Empirical results on a set of exploration games, challenging cooperative predator-prey tasks requiring significant coordination among agents, and StarCraft II micromanagement benchmarks show that UneVEn can solve tasks where other state-of-the-art MARL methods fail. Tarun Gupta 0002, Anuj Mahajan, Bei Peng 0001, Wendelin Böhmer, Shimon Whiteson |
ICML | 2 |
| 2021 | Tesseract: Tensorised Actors for Multi-Agent Reinforcement LearningabstractReinforcement Learning in large action spaces is a challenging problem. This is especially true for cooperative multi-agent reinforcement learning (MARL), which often requires tractable learning while respecting various constraints like communication budget and information about other agents. In this work, we focus on the fundamental hurdle affecting both value-based and policy-gradient approaches: an exponential blowup of the action space with the number of agents. For value-based methods, it poses challenges in accurately representing the optimal value function for value-based methods, thus inducing suboptimality. For policy gradient methods, it renders the critic ineffective and exacerbates the problem of the lagging critic. We show that from a learning theory perspective, both problems can be addressed by accurately representing the associated action-value function with a low-complexity hypothesis class. This requires accurately modelling the agent interactions in a sample efficient way. To this end, we propose a novel tensorised formulation of the Bellman equation. This gives rise to our method Tesseract, which utilises the view of Q-function seen as a tensor where the modes correspond to action spaces of different agents. Algorithms derived from Tesseract decompose the Q-tensor across the agents and utilise low-rank tensor approximations to model the agent interactions relevant to the task. We provide PAC analysis for Tesseract based algorithms and highlight their relevance to the class of rich observation MDPs. Empirical results in different domains confirm the gains in sample efficiency using Tesseract as supported by the theory. Anuj Mahajan, Mikayel Samvelyan, Viktor Makoviychuk, Animesh Garg, Jean Kossaifi, Shimon Whiteson, Yuke Zhu, Anima Anandkumar |
ICML | 1 |
| 2019 | VIREL: A Variational Inference Framework for Reinforcement LearningabstractApplying probabilistic models to reinforcement learning (RL) enables the uses of powerful optimisation tools such as variational inference in RL. However, existing inference frameworks and their algorithms pose significant challenges for learning optimal policies, e.g., the lack of mode capturing behaviour in pseudo-likelihood methods, difficulties learning deterministic policies in maximum entropy RL based approaches, and a lack of analysis when function approximators are used. We propose VIREL, a theoretically grounded probabilistic inference framework for RL that utilises a parametrised action-value function to summarise future dynamics of the underlying MDP, generalising existing approaches. VIREL also benefits from a mode-seeking form of KL divergence, the ability to learn deterministic optimal polices naturally from inference, and the ability to optimise value functions and policies in separate, iterative steps. In applying variational expectation-maximisation to VIREL, we thus show that the actor-critic algorithm can be reduced to expectation-maximisation, with policy improvement equivalent to an E-step and policy evaluation to an M-step. We then derive a family of actor-critic methods fromVIREL, including a scheme for adaptive exploration. Finally, we demonstrate that actor-critic algorithms from this family outperform state-of-the-art methods based on soft value functions in several domains. Matthew Fellows, Anuj Mahajan, Tim G. J. Rudner, Shimon Whiteson |
NeurIPS | 2 |
| 2019 | MAVEN: Multi-Agent Variational ExplorationabstractCentralised training with decentralised execution is an important setting for cooperative deep multi-agent reinforcement learning due to communication constraints during execution and computational tractability in training. In this paper, we analyse value-based methods that are known to have superior performance in complex environments. We specifically focus on QMIX, the current state-of-the-art in this domain. We show that the representation constraints on the joint action-values introduced by QMIX and similar methods lead to provably poor exploration and suboptimality. Furthermore, we propose a novel approach called MAVEN that hybridises value and policy-based methods by introducing a latent space for hierarchical control. The value-based agents condition their behaviour on the shared latent variable controlled by a hierarchical policy. This allows MAVEN to achieve committed, temporally extended exploration, which is key to solving complex multi-agent tasks. Our experimental results show that MAVEN achieves significant performance improvements on the challenging SMAC domain. Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, Shimon Whiteson |
NeurIPS | 1 |
| 2015 | Feature Selection for Short Text Classification using Wavelet Packet TransformabstractText classification tasks suffer from curse of dimensionality due to large feature space.Short text data further exacerbates the problem due to their sparse and noisy nature.Feature selection thus becomes an important step in improving the classification performance.In this paper, we propose a novel feature selection method using Wavelet Packet Transform.Wavelet Packet Transform (WPT) has been used widely in various fields due to its efficiency in encoding transient signals.We demonstrate how short text classification task can be benefited by feature selection using WPT due to their sparse nature.Our technique chooses the most discriminating features by computing inter-class distances in the transformed space.We experimented extensively with several short text datasets.Compared to well known techniques our approach reduces the feature space size and improves the overall classification performance significantly in all the datasets. Anuj Mahajan, Sharmistha Jat, Shourya Roy |
CoNLL | 1 |
| 2015 | Lifted Inference Rules With ConstraintsabstractLifted inference rules exploit symmetries for fast reasoning in statistical rela-tional models. Computational complexity of these rules is highly dependent onthe choice of the constraint language they operate on and therefore coming upwith the right kind of representation is critical to the success of lifted inference.In this paper, we propose a new constraint language, called setineq, which allowssubset, equality and inequality constraints, to represent substitutions over the vari-ables in the theory. Our constraint formulation is strictly more expressive thanexisting representations, yet easy to operate on. We reformulate the three mainlifting rules: decomposer, generalized binomial and the recently proposed singleoccurrence for MAP inference, to work with our constraint representation. Exper-iments on benchmark MLNs for exact and sampling based inference demonstratethe effectiveness of our approach over several other existing techniques. Happy Mittal, Anuj Mahajan, Vibhav Gogate, Parag Singla |
NIPS | 2 |
| 2008 | Mining Financial News for Major Events and Their Impacts on the MarketabstractIn this paper we have proposed a stock market analysis system that analyzes financial news items to identify and characterize major events that impact the market. The events have been identified using latent Dirichlet allocation (LDA) based topic extraction mechanism. These topics have been thereafter analyzed in conjunction with actual market data to understand their impact on the market. A prediction system has been proposed which can predict whether the stock market will fall or rise, based on news items. Anuj Mahajan, Lipika Dey, S. K. Mirajul Haque |
Web Intelligence | 1 |