Ameesh Shah

dblp:236/6361 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Multi-agent systems · 35% Optimization for machine learning · 20% Planning, search and constraint satisfaction · 20%
Theoretical computer science
2 papers
Automata and formal languages · 85% Algorithmic game theory and mechanism design · 15%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › minimax optimization
adversarial optimization
0.912025
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
0.912025
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › game theory
cooperative game
0.712023
Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.712023
Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › best-first search
a* search
0.412020
Learning Differentiable Programs with Admissible Neural Heuristics · NeurIPS 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search
0.412020
Learning Differentiable Programs with Admissible Neural Heuristics · NeurIPS 2020
Automata and formal languages
finite automata
0.412019
Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks · ICLR (Poster) 2019
Automata and formal languages
formal language representation
0.412019
Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks · ICLR (Poster) 2019
Automata and formal languages
recurrent neural networks
0.412019
Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks · ICLR (Poster) 2019
Algorithmic game theory and mechanism design › non-cooperative game
coordination games
0.212023
Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023

Methods — techniques the papers use, named apart from their topics

bellman backup operator · 1.3policy gradient · 0.9opponent shaping · 0.9neural heuristics · 0.9iterative deepening depth-first search · 0.9a* search · 0.9recurrent neural network · 0.4
YearPublicationVenuePosition
2025 Learning Symbolic Task Decompositions for Multi-Agent Teams
Ameesh Shah, Niklas Lauffer, Nikhil Pitta, Sanjit A. Seshia
AAMAS1
2025 Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
abstract
Adversarial optimization algorithms that explicitly search for flaws in agents' policies have been successfully applied to finding robust and diverse policies in the context of multi-agent learning. However, the success of adversarial optimization has been largely limited to zero-sum settings because its naive application in cooperative settings leads to a critical failure mode: agents are irrationally incentivized to *self-sabotage*, blocking the completion of tasks and halting further learning. To address this, we introduce *Rationality-preserving Policy Optimization (RPO)*, a formalism for adversarial optimization that avoids self-sabotage by ensuring agents remain *rational*—that is, their policies are optimal with respect to some possible partner policy. To solve RPO, we develop *Rational Policy Gradient (RPG)*, which trains agents to maximize their own reward in a modified version of the original game in which we use *opponent shaping* techniques to optimize the adversarial objective. RPG enables us to extend a variety of existing adversarial optimization algorithms that, no longer subject to the limitations of self-sabotage, can find adversarial examples, improve robustness and adaptability, and learn diverse policies. We empirically validate that our approach achieves strong performance in several popular cooperative and general-sum environments. Our project page can be found at https://rational-policy-gradient.github.io.
Niklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia, Stuart Russell 0001, Michael Dennis 0001
NeurIPS2
2023 Who Needs to Know? Minimal Knowledge for Optimal Coordination
abstract
To optimally coordinate with others in cooperative games, it is often crucial to have information about one’s collaborators: successful driving requires understanding which side of the road to drive on. However, not every feature of collaborators is strategically relevant: the fine-grained acceleration of drivers may be ignored while maintaining optimal coordination. We show that there is a well-defined dichotomy between strategically relevant and irrelevant information. Moreover, we show that, in dynamic games, this dichotomy has a compact representation that can be efficiently computed via a Bellman backup operator. We apply this algorithm to analyze the strategically relevant information for tasks in both a standard and a partially observable version of the Overcooked environment. Theoretical and empirical results show that our algorithms are significantly more efficient than baselines. Videos are available at https://minknowledge.github.io.
Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis 0001, Stuart Russell 0001
ICML2
2022 Learning Deterministic Finite Automata Decompositions from Examples and Demonstrations
Niklas Lauffer, Beyazit Yalcinkaya, Marcell Vazquez-Chanlatte, Ameesh Shah, Sanjit A. Seshia
FMCAD4
2020 Learning Differentiable Programs with Admissible Neural Heuristics
abstract
We study the problem of learning differentiable functions expressed as programs in a domain-specific language. Such programmatic models can offer benefits such as composability and interpretability; however, learning them requires optimizing over a combinatorial space of program "architectures". We frame this optimization problem as a search in a weighted graph whose paths encode top-down derivations of program syntax. Our key innovation is to view various classes of neural networks as continuous relaxations over the space of programs, which can then be used to complete any partial program. All the parameters of this relaxed program can be trained end-to-end, and the resulting training loss is an approximately admissible heuristic that can guide the combinatorial search. We instantiate our approach on top of the A* and Iterative Deepening Depth-First Search algorithms and use these algorithms to learn programmatic classifiers in three sequence classification tasks. Our experiments show that the algorithms outperform state-of-the-art methods for program learning, and that they discover programmatic classifiers that yield natural interpretations and achieve competitive accuracy.
Ameesh Shah, Eric Zhan, Jennifer J. Sun, Abhinav Verma 0001, Yisong Yue, Swarat Chaudhuri
NeurIPS1
2019 Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks
Joshua J. Michalenko, Ameesh Shah, Abhinav Verma 0001, Richard G. Baraniuk, Swarat Chaudhuri, Ankit B. Patel
ICLR (Poster)2