EDBT 2026 Demo / reviewers in the wild / expert
Micah Carroll
dblp:250/9080 · also Micah D. Carroll
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 44% Multi-agent systems · 14% Optimization for machine learning · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.0 | 2 | 2025 | AI Alignment with Changing and Influenceable Reward Functions · ICML 2024 On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback · ICLR 2025 |
Machine learning › Optimization for machine learning › minimax optimization
adversarial optimization |
0.9 | 1 | 2025 | Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning |
0.9 | 1 | 2025 | Robust and Diverse Multi-Agent Learning via Rational Policy Gradient · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback · ICLR 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.8 | 1 | 2024 | AI Alignment with Changing and Influenceable Reward Functions · ICML 2024 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
cooperative game |
0.7 | 1 | 2023 | Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023 |
Machine learning › Reinforcement learning
partially observable reinforcement learning |
0.7 | 1 | 2023 | Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023 |
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning |
0.6 | 1 | 2022 | Uni[MASK]: Unified Inference in Sequential Decision Problems · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction |
0.6 | 1 | 2022 | Uni[MASK]: Unified Inference in Sequential Decision Problems · NeurIPS 2022 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.6 | 1 | 2022 | Uni[MASK]: Unified Inference in Sequential Decision Problems · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.6 | 1 | 2022 | Uni[MASK]: Unified Inference in Sequential Decision Problems · NeurIPS 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
human-AI collaboration |
0.4 | 1 | 2019 | On the Utility of Learning about Humans for Human-AI Coordination · NeurIPS 2019 |
Machine learning › Optimization for machine learning › hyperparameter optimization
population-based training |
0.4 | 1 | 2019 | On the Utility of Learning about Humans for Human-AI Coordination · NeurIPS 2019 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
0.4 | 1 | 2019 | On the Utility of Learning about Humans for Human-AI Coordination · NeurIPS 2019 |
Algorithmic game theory and mechanism design › non-cooperative game
coordination games |
0.2 | 1 | 2023 | Who Needs to Know? Minimal Knowledge for Optimal Coordination · ICML 2023 |
Machine learning › Trustworthy machine learning
fairness |
0.2 | 1 | 2022 | Estimating and Penalizing Induced Preference Shifts in Recommender Systems · ICML 2022 |
Natural language and speech › Language models and text generation
masked language modeling |
0.2 | 1 | 2022 | Uni[MASK]: Unified Inference in Sequential Decision Problems · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
bellman backup operator · 1.3simulation · 1.1predictive user modeling · 1.1simulated user feedback · 0.9reinforcement learning · 0.9policy gradient · 0.9opponent shaping · 0.9LLM-as-judge · 0.9markov decision process · 0.8fine-tuning · 0.6user study · 0.4self-play · 0.4population-based training · 0.4planning algorithm · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Targeted Manipulation and Deception when Optimizing LLMs for User FeedbackabstractAs LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative or deceptive tactics to obtain positive feedback from users who are vulnerable to such strategies. We study this phenomenon by training LLMs with Reinforcement Learning with simulated user feedback in environments of practical LLM usage. In our settings, we find that: 1) Extreme forms of "feedback gaming" such as manipulation and deception are learned reliably; 2) Even if only 2% of users are vulnerable to manipulative strategies, LLMs learn to identify and target them while behaving appropriately with other users, making such behaviors harder to detect; 3) To mitigate this issue, it may seem promising to leverage continued safety training or LLM-as-judges during training to filter problematic outputs. Instead, we found that while such approaches help in some of our settings, they backfire in others, sometimes even leading to subtler manipulative behaviors. We hope our results can serve as a case study which highlights the risks of using gameable feedback sources -- such as user feedback -- as a target for RL. Our code is publicly available. Warning: some of our examples may be upsetting. Marcus Williams, Micah Carroll, Adhyyan Narang, Constantin Weisser, Brendan Murphy, Anca D. Dragan |
ICLR | 2 |
| 2025 | Robust and Diverse Multi-Agent Learning via Rational Policy GradientabstractAdversarial optimization algorithms that explicitly search for flaws in agents' policies have been successfully applied to finding robust and diverse policies in the context of multi-agent learning. However, the success of adversarial optimization has been largely limited to zero-sum settings because its naive application in cooperative settings leads to a critical failure mode: agents are irrationally incentivized to *self-sabotage*, blocking the completion of tasks and halting further learning. To address this, we introduce *Rationality-preserving Policy Optimization (RPO)*, a formalism for adversarial optimization that avoids self-sabotage by ensuring agents remain *rational*—that is, their policies are optimal with respect to some possible partner policy. To solve RPO, we develop *Rational Policy Gradient (RPG)*, which trains agents to maximize their own reward in a modified version of the original game in which we use *opponent shaping* techniques to optimize the adversarial objective. RPG enables us to extend a variety of existing adversarial optimization algorithms that, no longer subject to the limitations of self-sabotage, can find adversarial examples, improve robustness and adaptability, and learn diverse policies. We empirically validate that our approach achieves strong performance in several popular cooperative and general-sum environments. Our project page can be found at https://rational-policy-gradient.github.io. Niklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia, Stuart Russell 0001, Michael Dennis 0001 |
NeurIPS | 3 |
| 2024 | AI Alignment with Changing and Influenceable Reward FunctionsabstractExisting AI alignment approaches assume that preferences are static, which is unrealistic: our preferences change, and may even be influenced by our interactions with AI systems themselves. To clarify the consequences of incorrectly assuming static preferences, we introduce Dynamic Reward Markov Decision Processes (DR-MDPs), which explicitly model preference changes and the AI’s influence on them. We show that despite its convenience, the static-preference assumption may undermine the soundness of existing alignment techniques, leading them to implicitly reward AI systems for influencing user preferences in ways users may not truly want. We then explore potential solutions. First, we offer a unifying perspective on how an agent’s optimization horizon may partially help reduce undesirable AI influence. Then, we formalize different notions of AI alignment that account for preference change from the outset. Comparing the strengths and limitations of 8 such notions of alignment, we find that they all either err towards causing undesirable AI influence, or are overly risk-averse, suggesting that a straightforward solution to the problems of changing preferences may not exist. As there is no avoiding grappling with changing preferences in real-world settings, this makes it all the more important to handle these issues with care, balancing risks and capabilities. We hope our work can provide conceptual clarity and constitute a first step towards AI alignment practices which explicitly account for (and contend with) the changing and influenceable nature of human preferences. Micah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 0001, Anca D. Dragan |
ICML | 1 |
| 2023 | Who Needs to Know? Minimal Knowledge for Optimal CoordinationabstractTo optimally coordinate with others in cooperative games, it is often crucial to have information about one’s collaborators: successful driving requires understanding which side of the road to drive on. However, not every feature of collaborators is strategically relevant: the fine-grained acceleration of drivers may be ignored while maintaining optimal coordination. We show that there is a well-defined dichotomy between strategically relevant and irrelevant information. Moreover, we show that, in dynamic games, this dichotomy has a compact representation that can be efficiently computed via a Bellman backup operator. We apply this algorithm to analyze the strategically relevant information for tasks in both a standard and a partially observable version of the Overcooked environment. Theoretical and empirical results show that our algorithms are significantly more efficient than baselines. Videos are available at https://minknowledge.github.io. Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis 0001, Stuart Russell 0001 |
ICML | 3 |
| 2022 | Estimating and Penalizing Induced Preference Shifts in Recommender SystemsabstractThe content that a recommender system (RS) shows to users influences them. Therefore, when choosing a recommender to deploy, one is implicitly also choosing to induce specific internal states in users. Even more, systems trained via long-horizon optimization will have direct incentives to manipulate users, e.g. shift their preferences so they are easier to satisfy. We focus on induced preference shifts in users. We argue that {–} before deployment {–} system designers should: estimate the shifts a recommender would induce; evaluate whether such shifts would be undesirable; and perhaps even actively optimize to avoid problematic shifts. These steps involve two challenging ingredients: estimation requires anticipating how hypothetical policies would influence user preferences if deployed {–} we do this by using historical user interaction data to train a predictive user model which implicitly contains their preference dynamics; evaluation and optimization additionally require metrics to assess whether such influences are manipulative or otherwise unwanted {–} we use the notion of "safe shifts", that define a trust region within which behavior is safe: for instance, the natural way in which users would shift without interference from the system could be deemed "safe". In simulated experiments, we show that our learned preference dynamics model is effective in estimating user preferences and how they would respond to new recommenders. Additionally, we show that recommenders that optimize for staying in the trust region can avoid manipulative behaviors while still generating engagement. Micah Carroll, Anca D. Dragan, Stuart Russell 0001, Dylan Hadfield-Menell |
ICML | 1 |
| 2022 | Uni[MASK]: Unified Inference in Sequential Decision ProblemsabstractRandomly masking and predicting word tokens has been a successful approach in pre-training language models for a variety of downstream tasks. In this work, we observe that the same idea also applies naturally to sequential decision making, where many well-studied tasks like behavior cloning, offline RL, inverse dynamics, and waypoint conditioning correspond to different sequence maskings over a sequence of states, actions, and returns. We introduce the UniMASK framework, which provides a unified way to specify models which can be trained on many different sequential decision making tasks. We show that a single UniMASK model is often capable of carrying out many tasks with performance similar to or better than single-task models. Additionally, after fine-tuning, our UniMASK models consistently outperform comparable single-task models. Micah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu, Mingfei Sun 0001, David Bignell, Stephanie Milani, Katja Hofmann, Matthew J. Hausknecht, Anca D. Dragan, Sam Devlin |
NeurIPS | 1 |
| 2021 | Estimating and Penalizing Preference Shift in Recommender SystemsabstractRecommender systems trained via long-horizon optimization (e.g., reinforcement learning) will have incentives to actively manipulate user preferences through the recommended content. While some work has argued for making systems myopic to avoid this issue, even such systems can induce systematic undesirable preference shifts. Thus, rather than artificially stifling the capabilities of the system, in this work we explore how we can make capable systems that explicitly avoid undesirable shifts. We advocate for (1) estimating the preference shifts that would be induced by recommender system policies, and (2) explicitly characterizing what unwanted shifts are and assessing before deployment whether such policies will produce them – ideally even actively optimizing to avoid them. These steps involve two challenging ingredients: (1) requires the ability to anticipate how hypothetical policies would influence user preferences if deployed; instead, (2) requires metrics to assess whether such influences are manipulative or otherwise unwanted. We study how to do (1) from historical user interaction data by building a user predictive model that implicitly contains their preference dynamics; to address (2), we introduce the notion of a “safe policy”, which defines a trust region within which behavior is believed to be safe. We show that recommender systems that optimize for staying in the trust region avoid manipulative behaviors (e.g., changing preferences in ways that make users more predictable), while still generating engagement. Micah Carroll, Dylan Hadfield-Menell, Stuart Russell 0001, Anca D. Dragan |
RecSys | 1 |
| 2019 | On the Utility of Learning about Humans for Human-AI CoordinationabstractWhile we would like agents that can coordinate with humans, current algorithms such as self-play and population-based training create agents that can coordinate with themselves. Agents that assume their partner to be optimal or similar to them can converge to coordination protocols that fail to understand and be understood by humans. To demonstrate this, we introduce a simple environment that requires challenging coordination, based on the popular game Overcooked, and learn a simple model that mimics human play. We evaluate the performance of agents trained via self-play and population-based training. These agents perform very well when paired with themselves, but when paired with our human model, they are significantly worse than agents designed to play with the human model. An experiment with a planning algorithm yields the same conclusion, though only when the human-aware planner is given the exact human model that it is playing with. A user study with real humans shows this pattern as well, though less strongly. Qualitatively, we find that the gains come from having the agent adapt to the human's gameplay. Given this result, we suggest several approaches for designing agents that learn about humans in order to better coordinate with them. Code is available at https://github.com/HumanCompatibleAI/overcooked_ai. Micah Carroll, Rohin Shah, Mark K. Ho, Thomas L. Griffiths 0001, Sanjit A. Seshia, Pieter Abbeel, Anca D. Dragan |
NeurIPS | 1 |