EDBT 2026 Demo / reviewers in the wild / expert
Niklas Lauffer
dblp:321/0666
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 34% Multi-agent systems · 24% Optimization for machine learning · 14% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › minimax optimization
adversarial optimization |
0.9 | 1 | 2025 | Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning |
0.9 | 1 | 2025 | Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025 |
Machine learning › Learning theory
finite state machine |
0.8 | 1 | 2024 | Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning |
0.8 | 1 | 2024 | Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
goal representation |
0.8 | 1 | 2024 | Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › temporal planning
temporal goal specification |
0.8 | 1 | 2024 | Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
cooperative game |
0.7 | 1 | 2023 | Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023 |
Machine learning › Reinforcement learning
partially observable reinforcement learning |
0.7 | 1 | 2023 | Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023 |
Machine learning › Graph learning › graph representation learning
graph neural network representation learning |
0.2 | 1 | 2024 | Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024 |
Algorithmic game theory and mechanism design › non-cooperative game
coordination games |
0.2 | 1 | 2023 | Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
bellman backup operator · 1.3policy gradient · 0.9opponent shaping · 0.9pre-training · 0.8graph neural network · 0.8automata composition · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Symbolic Task Decompositions for Multi-Agent Teams
Ameesh Shah, Niklas Lauffer, Nikhil Pitta, Sanjit A. Seshia |
AAMAS | 2 |
| 2025 | Robust and Diverse Multi-Agent Learning via Rational Policy GradientabstractAdversarial optimization algorithms that explicitly search for flaws in agents' policies have been successfully applied to finding robust and diverse policies in the context of multi-agent learning. However, the success of adversarial optimization has been largely limited to zero-sum settings because its naive application in cooperative settings leads to a critical failure mode: agents are irrationally incentivized to *self-sabotage*, blocking the completion of tasks and halting further learning. To address this, we introduce *Rationality-preserving Policy Optimization (RPO)*, a formalism for adversarial optimization that avoids self-sabotage by ensuring agents remain *rational*—that is, their policies are optimal with respect to some possible partner policy. To solve RPO, we develop *Rational Policy Gradient (RPG)*, which trains agents to maximize their own reward in a modified version of the original game in which we use *opponent shaping* techniques to optimize the adversarial objective. RPG enables us to extend a variety of existing adversarial optimization algorithms that, no longer subject to the limitations of self-sabotage, can find adversarial examples, improve robustness and adaptability, and learn diverse policies. We empirically validate that our approach achieves strong performance in several popular cooperative and general-sum environments. Our project page can be found at https://rational-policy-gradient.github.io. Niklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia, Stuart Russell 0001, Michael Dennis 0001 |
NeurIPS | 1 |
| 2024 | Compositional Automata Embeddings for Goal-Conditioned Reinforcement LearningabstractGoal-conditioned reinforcement learning is a powerful way to control an AI agent's behavior at runtime. That said, popular goal representations, e.g., target states or natural language, are either limited to Markovian tasks or rely on ambiguous task semantics. We propose representing temporal goals using compositions of deterministic finite automata (cDFAs) and use cDFAs to guide RL agents. cDFAs balance the need for formal temporal semantics with ease of interpretation: if one can understand a flow chart, one can understand a cDFA. On the other hand, cDFAs form a countably infinite concept class with Boolean semantics, and subtle changes to the automaton can result in very different tasks, making them difficult to condition agent behavior on. To address this, we observe that all paths through a DFA correspond to a series of reach-avoid tasks and propose pre-training graph neural network embeddings on "reach-avoid derived" DFAs. Through empirical evaluation, we demonstrate that the proposed pre-training method enables zero-shot generalization to various cDFA task classes and accelerated policy specialization without the myopic suboptimality of hierarchical methods. Beyazit Yalcinkaya, Niklas Lauffer, Marcell Vazquez-Chanlatte, Sanjit A. Seshia |
NeurIPS | 2 |
| 2023 | Who Needs to Know? Minimal Knowledge for Optimal CoordinationabstractTo optimally coordinate with others in cooperative games, it is often crucial to have information about one’s collaborators: successful driving requires understanding which side of the road to drive on. However, not every feature of collaborators is strategically relevant: the fine-grained acceleration of drivers may be ignored while maintaining optimal coordination. We show that there is a well-defined dichotomy between strategically relevant and irrelevant information. Moreover, we show that, in dynamic games, this dichotomy has a compact representation that can be efficiently computed via a Bellman backup operator. We apply this algorithm to analyze the strategically relevant information for tasks in both a standard and a partially observable version of the Overcooked environment. Theoretical and empirical results show that our algorithms are significantly more efficient than baselines. Videos are available at https://minknowledge.github.io. Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis 0001, Stuart Russell 0001 |
ICML | 1 |
| 2022 | Learning Deterministic Finite Automata Decompositions from Examples and Demonstrations
Niklas Lauffer, Beyazit Yalcinkaya, Marcell Vazquez-Chanlatte, Ameesh Shah, Sanjit A. Seshia |
FMCAD | 1 |