Niklas Lauffer

dblp:321/0666 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 34% Multi-agent systems · 24% Optimization for machine learning · 14%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › minimax optimization
adversarial optimization
0.912025
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
0.912025
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025
Machine learning › Learning theory
finite state machine
0.812024
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.812024
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
goal representation
0.812024
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › temporal planning
temporal goal specification
0.812024
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems › game theory
cooperative game
0.712023
Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.712023
Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023
Machine learning › Graph learning › graph representation learning
graph neural network representation learning
0.212024
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning · NeurIPS 2024
Algorithmic game theory and mechanism design › non-cooperative game
coordination games
0.212023
Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023

Methods — techniques the papers use, named apart from their topics

bellman backup operator · 1.3policy gradient · 0.9opponent shaping · 0.9pre-training · 0.8graph neural network · 0.8automata composition · 0.8
YearPublicationVenuePosition
2025 Learning Symbolic Task Decompositions for Multi-Agent Teams
Ameesh Shah, Niklas Lauffer, Nikhil Pitta, Sanjit A. Seshia
AAMAS2
2025 Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
abstract
Adversarial optimization algorithms that explicitly search for flaws in agents' policies have been successfully applied to finding robust and diverse policies in the context of multi-agent learning. However, the success of adversarial optimization has been largely limited to zero-sum settings because its naive application in cooperative settings leads to a critical failure mode: agents are irrationally incentivized to *self-sabotage*, blocking the completion of tasks and halting further learning. To address this, we introduce *Rationality-preserving Policy Optimization (RPO)*, a formalism for adversarial optimization that avoids self-sabotage by ensuring agents remain *rational*—that is, their policies are optimal with respect to some possible partner policy. To solve RPO, we develop *Rational Policy Gradient (RPG)*, which trains agents to maximize their own reward in a modified version of the original game in which we use *opponent shaping* techniques to optimize the adversarial objective. RPG enables us to extend a variety of existing adversarial optimization algorithms that, no longer subject to the limitations of self-sabotage, can find adversarial examples, improve robustness and adaptability, and learn diverse policies. We empirically validate that our approach achieves strong performance in several popular cooperative and general-sum environments. Our project page can be found at https://rational-policy-gradient.github.io.
Niklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia, Stuart Russell 0001, Michael Dennis 0001
NeurIPS1
2024 Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
abstract
Goal-conditioned reinforcement learning is a powerful way to control an AI agent's behavior at runtime. That said, popular goal representations, e.g., target states or natural language, are either limited to Markovian tasks or rely on ambiguous task semantics. We propose representing temporal goals using compositions of deterministic finite automata (cDFAs) and use cDFAs to guide RL agents. cDFAs balance the need for formal temporal semantics with ease of interpretation: if one can understand a flow chart, one can understand a cDFA. On the other hand, cDFAs form a countably infinite concept class with Boolean semantics, and subtle changes to the automaton can result in very different tasks, making them difficult to condition agent behavior on. To address this, we observe that all paths through a DFA correspond to a series of reach-avoid tasks and propose pre-training graph neural network embeddings on "reach-avoid derived" DFAs. Through empirical evaluation, we demonstrate that the proposed pre-training method enables zero-shot generalization to various cDFA task classes and accelerated policy specialization without the myopic suboptimality of hierarchical methods.
Beyazit Yalcinkaya, Niklas Lauffer, Marcell Vazquez-Chanlatte, Sanjit A. Seshia
NeurIPS2
2023 Who Needs to Know? Minimal Knowledge for Optimal Coordination
abstract
To optimally coordinate with others in cooperative games, it is often crucial to have information about one’s collaborators: successful driving requires understanding which side of the road to drive on. However, not every feature of collaborators is strategically relevant: the fine-grained acceleration of drivers may be ignored while maintaining optimal coordination. We show that there is a well-defined dichotomy between strategically relevant and irrelevant information. Moreover, we show that, in dynamic games, this dichotomy has a compact representation that can be efficiently computed via a Bellman backup operator. We apply this algorithm to analyze the strategically relevant information for tasks in both a standard and a partially observable version of the Overcooked environment. Theoretical and empirical results show that our algorithms are significantly more efficient than baselines. Videos are available at https://minknowledge.github.io.
Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis 0001, Stuart Russell 0001
ICML1
2022 Learning Deterministic Finite Automata Decompositions from Examples and Demonstrations
Niklas Lauffer, Beyazit Yalcinkaya, Marcell Vazquez-Chanlatte, Ameesh Shah, Sanjit A. Seshia
FMCAD1