EDBT 2026 Demo / reviewers in the wild / expert
Mark K. Ho
dblp:191/6682
· DBLP profile ↗
36ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0002-1454-4768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 8 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Integration of Language and Experience via the Instructed Bandit Task
Ellen Su, Mark K. Ho, Todd M. Gureckis |
CogSci | 2 |
| 2025 | Learning about Inductive Potential from Generic Statements
Marianna Y. Zhang, Sarah-Jane Leslie, Marjorie Rhodes, Mark K. Ho |
CogSci | 4 |
| 2025 | Estimating cognitive biases with attention-aware inverse planningabstractPeople's goal-directed behaviors are influenced by their cognitive biases, and autonomous systems that interact with people should be aware of this. For example, people's attention to objects in their environment will be biased in a way that systematically affects how they perform everyday tasks such as driving to work. Here, building on recent work in computational cognitive science, we formally articulate the \textit{attention-aware inverse planning problem}, in which the goal is to estimate a person's attentional biases from their actions. We demonstrate how attention-aware inverse planning systematically differs from standard inverse reinforcement learning and how cognitive biases can be inferred from behavior. Finally, we present an approach to attention-aware inverse planning that combines deep reinforcement learning with computational cognitive modeling. We use this approach to infer the attentional strategies of RL agents in real-life driving scenarios selected from the Waymo Open Dataset, demonstrating the scalability of estimating cognitive biases with attention-aware inverse planning. Sounak Banerjee 0002, Daphne Cornelisse, Deepak Edakkattil Gopinath, Emily S. Sumner, Jonathan A. DeCastro, Guy Rosman, Eugene Vinitsky, Mark K. Ho |
NeurIPS | 8 |
| 2024 | Structurally Guided Task Decomposition in Spatial Navigation Tasks (Student Abstract)abstractHow are people able to plan so efficiently despite limited cognitive resources? We aimed to answer this question by extending an existing model of human task decomposition that can explain a wide range of simple planning problems by adding structure information to the task to facilitate planning in more complex tasks. The extended model was then applied to a more complex planning domain of spatial navigation. Our results suggest that our framework can correctly predict the navigation strategies of the majority of the participants in an online experiment. Ruiqi He, Carlos G. Correa, Thomas L. Griffiths 0001, Mark K. Ho |
AAAI | 4 |
| 2024 | Modeling Cognitive Strategies in Teaching: Integrating Theory of Mind and Heuristics
Sevan K. Harootonian, Yael Niv, Thomas L. Griffiths 0001, Mark K. Ho |
CogSci | 4 |
| 2024 | Common sense reasoning about credibility
Peiyao Hu, Mark K. Ho |
CogSci | 2 |
| 2024 | Teaching Functions with Gaussian Process Regression
Maya Malaviya, Mark K. Ho |
CogSci | 2 |
| 2024 | Investigating Flexible Role Binding in AI Agents
Brian Pennisi, Todd M. Gureckis, Rheza Budiono, Mark K. Ho |
CogSci | 4 |
| 2024 | Concept Alignment as a Prerequisite for Value Alignment
Sunayana Rane, Mark K. Ho, Ilia Sucholutsky, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2023 | Diagnosis, Feedback, Adaptation: A Human-in-the-Loop Framework for Test-Time Policy AdaptationabstractPolicies often fail at test-time due to distribution shifts—changes in the state and reward that occur when an end user deploys the policy in environments different from those seen in training. Data augmentation can help models be more robust to such shifts by varying specific concepts in the state, e.g. object color, that are task-irrelevant and should not impact desired actions. However, designers training the agent don’t often know which concepts are irrelevant a priori. We propose a human-in-the-loop framework to leverage feedback from the end user to quickly identify and augment task-irrelevant visual state concepts. Our framework generates counterfactual demonstrations that allow users to quickly isolate shifted state concepts and identify if they should not impact the desired task, and can therefore be augmented using existing actions. We present experiments validating our full pipeline on discrete and continuous control tasks with real human users. Our method better enables users to (1) understand agent failure, (2) improve sample efficiency of demonstrations required for finetuning, and (3) adapt the agent to their desired reward. Andi Peng, Aviv Netanyahu, Mark K. Ho, Tianmin Shu, Andreea Bobu, Julie A. Shah, Pulkit Agrawal 0001 |
ICML | 3 |
| 2023 | Humans decompose tasks by trading off utility and computational costabstractHuman behavior emerges from planning over elaborate decompositions of tasks into goals, subgoals, and low-level actions. How are these decompositions created and used? Here, we propose and evaluate a normative framework for task decomposition based on the simple idea that people decompose tasks to reduce the overall cost of planning while maintaining task performance. Analyzing 11,117 distinct graph-structured planning tasks, we find that our framework justifies several existing heuristics for task decomposition and makes predictions that can be distinguished from two alternative normative accounts. We report a behavioral study of task decomposition (N = 806) that uses 30 randomly sampled graphs, a larger and more diverse set than that of any previous behavioral study on this topic. We find that human responses are more consistent with our framework for task decomposition than alternative normative accounts and are most consistent with a heuristic-betweenness centrality-that is justified by our approach. Taken together, our results suggest the computational cost of planning is a key principle guiding the intelligent structuring of goal-directed behavior. Carlos G. Correa, Mark K. Ho, Frederick Callaway, Nathaniel D. Daw, Thomas L. Griffiths 0001 |
PLoS Comput. Biol. | 2 |
| 2022 | On the Expressivity of Markov Reward (Extended Abstract)abstractReward is the driving force for reinforcement-learning agents. We here set out to understand the expressivity of Markov reward as a way to capture tasks that we would want an agent to perform. We frame this study around three new abstract notions of "task": (1) a set of acceptable behaviors, (2) a partial ordering over behaviors, or (3) a partial ordering over trajectories. Our main results prove that while reward can express many of these tasks, there exist instances of each task type that no Markov reward function can capture. We then provide a set of polynomial-time algorithms that construct a Markov reward function that allows an agent to perform each task type, and correctly determine when no such reward function exists. David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho, Michael L. Littman, Doina Precup, Satinder Singh 0001 |
IJCAI | 4 |
| 2022 | How to talk so AI will learn: Instructions, descriptions, and autonomyabstractFrom the earliest years of our lives, humans use language to express our beliefs and desires. Being able to talk to artificial agents about our preferences would thus fulfill a central goal of value alignment. Yet today, we lack computational models explaining such language use. To address this challenge, we formalize learning from language in a contextual bandit setting and ask how a human might communicate preferences over behaviors. We study two distinct types of language: instructions, which provide information about the desired policy, and descriptions, which provide information about the reward function. We show that the agent's degree of autonomy determines which form of language is optimal: instructions are better in low-autonomy settings, but descriptions are better when the agent will need to act independently. We then define a pragmatic listener agent that robustly infers the speaker's reward function by reasoning about how the speaker expresses themselves. We validate our models with a behavioral experiment, demonstrating that (1) our speaker model predicts human behavior, and (2) our pragmatic listener successfully recovers humans' reward functions. Finally, we show that this form of social learning can integrate with and reduce regret in traditional reinforcement learning. We hope these insights facilitate a shift from developing agents that obey language to agents that learn from it. Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths 0001, Dylan Hadfield-Menell |
NeurIPS | 3 |
| 2021 | Learning Rewards From Linguistic FeedbackabstractWe explore unconstrained natural language feedback as a learning signal for artificial agents. Humans use rich and varied language to teach, yet most prior work on interactive learning from language assumes a particular form of input (e.g., commands). We propose a general framework which does not make this assumption, instead using aspect-based sentiment analysis to decompose feedback into sentiment over the features of a Markov decision process. We then infer the teacher's reward function by regressing the sentiment on the features, an analogue of inverse reinforcement learning. To evaluate our approach, we first collect a corpus of teaching behavior in a cooperative task where both teacher and learner are human. We implement three artificial learners: sentiment-based "literal" and "pragmatic" models, and an inference network trained end-to-end to predict rewards. We then re-run our initial experiment, pairing human teachers with these artificial learners. All three models successfully learn from interactive human feedback. The inference network approaches the performance of the "literal" sentiment model, while the "pragmatic" model nears human performance. Our work provides insight into the information structure of naturalistic linguistic feedback as well as methods to leverage it for reinforcement learning. Theodore R. Sumers, Mark K. Ho, Robert D. Hawkins, Karthik Narasimhan, Thomas L. Griffiths 0001 |
AAAI | 2 |
| 2021 | Using Machine Teaching to Investigate Human Assumptions when Teaching Reinforcement Learners
Yun-Shiuan Chuang, Xuezhou Zhang, Yuzhe Ma, Mark K. Ho, Joseph L. Austerweil, Jerry Zhu |
CogSci | 4 |
| 2021 | Extending rational models of communication from beliefs to actions
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths 0001 |
CogSci | 3 |
| 2021 | Specialization and selective social attention establishes the balance between individual and social learning
Charley M. Wu, Mark K. Ho, Benjamin Kahl, Christina Leuker, Björn Meder, Ralf H. J. M. Kurvers |
CogSci | 2 |
| 2021 | On the Expressivity of Markov RewardabstractReward is the driving force for reinforcement-learning agents. This paper is dedicated to understanding the expressivity of reward as a way to capture tasks that we would want an agent to perform. We frame this study around three new abstract notions of “task” that might be desirable: (1) a set of acceptable behaviors, (2) a partial ordering over behaviors, or (3) a partial ordering over trajectories. Our main results prove that while reward can express many of these tasks, there exist instances of each task type that no Markov reward function can capture. We then provide a set of polynomial-time algorithms that construct a Markov reward function that allows an agent to optimize tasks of each of these three types, and correctly determine when no such reward function exists. We conclude with an empirical study that corroborates and illustrates our theoretical findings. David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho, Michael L. Littman, Doina Precup, Satinder Singh 0001 |
NeurIPS | 4 |
| 2020 | People Do Not Just Plan, They Plan to Plan
Mark K. Ho, David Abel, Jonathan D. Cohen 0003, Michael L. Littman, Thomas L. Griffiths 0001 |
AAAI | 1 |
| 2020 | Resource-rational Task Decomposition to Minimize Planning Costs
Carlos G. Correa, Mark K. Ho, Frederick Callaway, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2020 | Punishment: Incentive or Communication?
Arunima Sarin, Mark K. Ho, Justin Martin, Fiery Cushman |
CogSci | 2 |
| 2020 | Show or Tell? Demonstration is More Robust to Changes in Shared Perception than Explanation
Theodore R. Sumers, Mark K. Ho, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2020 | Cognition, Collectives, and Human Culture
Charley M. Wu, Natalia Vélez, Mark K. Ho, Robert L. Goldstone |
CogSci | 3 |
| 2020 | Teaching a Robot Tasks of Arbitrary Complexity via Human FeedbackabstractThis paper addresses the problem of training a robot to carry out temporal tasks of arbitrary complexity via evaluative human feedback that can be inaccurate. A key idea explored in our work is a kind of curriculum learning---training the robot to master simple tasks and then building up to more complex tasks. We show how a training procedure, using knowledge of the formal task representation, can decompose and train any task efficiently in the size of its representation. We further provide a set of experiments that support the claim that non-expert human trainers can decompose tasks in a way that is consistent with our theoretical results, with more than half of participants successfully training all of our experimental missions. We compared our algorithm with existing approaches and our experimental results suggest that our method outperforms alternatives, especially when feedback contains mistakes. Carl Trimbach, Jun Ki Lee, Mark K. Ho, Michael L. Littman |
HRI | 4 |
| 2019 | Compositional subgoal representations
Carlos G. Correa, Frederick Callaway, Mark K. Ho, Thomas L. Griffiths 0001 |
CogSci | 3 |
| 2019 | The Computational Structure of Unintentional Meaning
Mark K. Ho, Joanna Korman, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2019 | On the Utility of Learning about Humans for Human-AI CoordinationabstractWhile we would like agents that can coordinate with humans, current algorithms such as self-play and population-based training create agents that can coordinate with themselves. Agents that assume their partner to be optimal or similar to them can converge to coordination protocols that fail to understand and be understood by humans. To demonstrate this, we introduce a simple environment that requires challenging coordination, based on the popular game Overcooked, and learn a simple model that mimics human play. We evaluate the performance of agents trained via self-play and population-based training. These agents perform very well when paired with themselves, but when paired with our human model, they are significantly worse than agents designed to play with the human model. An experiment with a planning algorithm yields the same conclusion, though only when the human-aware planner is given the exact human model that it is playing with. A user study with real humans shows this pattern as well, though less strongly. Qualitatively, we find that the gains come from having the agent adapt to the human's gameplay. Given this result, we suggest several approaches for designing agents that learn about humans in order to better coordinate with them. Code is available at https://github.com/HumanCompatibleAI/overcooked_ai. Micah Carroll, Rohin Shah, Mark K. Ho, Thomas L. Griffiths 0001, Sanjit A. Seshia, Pieter Abbeel, Anca D. Dragan |
NeurIPS | 3 |
| 2018 | Effectively Learning from Pedagogical Demonstrations
Mark K. Ho, Michael L. Littman, Fiery Cushman, Joseph L. Austerweil |
CogSci | 1 |
| 2018 | Learning Task Specifications from DemonstrationsabstractReal-world applications often naturally decompose into several sub-tasks. In many settings (e.g., robotics) demonstrations provide a natural way to specify the sub-tasks. However, most methods for learning from demonstrations either do not provide guarantees that the artifacts learned for the sub-tasks can be safely recombined or limit the types of composition available. Motivated by this deficit, we consider the problem of inferring Boolean non-Markovian rewards (also known as logical trace properties or specifications) from demonstrations provided by an agent operating in an uncertain, stochastic environment. Crucially, specifications admit well-defined composition rules that are typically easy to interpret. In this paper, we formulate the specification inference task as a maximum a posteriori (MAP) probability inference problem, apply the principle of maximum entropy to derive an analytic demonstration likelihood model and give an efficient approach to search for the most likely specification in a large candidate pool of specifications. In our experiments, we demonstrate how learning specifications can help avoid common problems that often arise due to ad-hoc reward composition. Marcell Vazquez-Chanlatte, Susmit Jha, Ashish Tiwari 0001, Mark K. Ho, Sanjit A. Seshia |
NeurIPS | 4 |
| 2017 | Teaching by Intervention: Working Backwards, Undoing Mistakes, or Correcting Mistakes?
Mark K. Ho, Michael L. Littman, Joseph L. Austerweil |
CogSci | 1 |
| 2017 | Interactive Learning from Policy-Dependent Human FeedbackabstractThis paper investigates the problem of interactively learning behaviors communicated by a human teacher using positive and negative feedback. Much previous work on this problem has made the assumption that people provide feedback for decisions that is dependent on the behavior they are teaching and is independent from the learner’s current policy. We present empirical results that show this assumption to be false—whether human trainers give a positive or negative feedback for a decision is influenced by the learner’s current policy. Based on this insight, we introduce Convergent Actor-Critic by Humans (COACH), an algorithm for learning from policy-dependent feedback that converges to a local optimum. Finally, we demonstrate that COACH can successfully learn multiple behaviors on a physical robot. James MacGlashan, Mark K. Ho, Robert Tyler Loftin, Bei Peng 0001, David L. Roberts 0001, Matthew E. Taylor, Michael L. Littman |
ICML | 2 |
| 2016 | Feature-based Joint Planning and Norm Learning in Collaborative Games
Mark K. Ho, James MacGlashan, Amy Greenwald, Michael L. Littman, Elizabeth Hilliard, Carl Trimbach, Stephen Brawner, Josh Tenenbaum, Max Kleiman-Weiner, Joseph L. Austerweil |
CogSci | 1 |
| 2016 | Coordinate to cooperate or compete: Abstract goals and joint intentions in social interaction
Max Kleiman-Weiner, Mark K. Ho, Joseph L. Austerweil, Michael L. Littman, Josh Tenenbaum |
CogSci | 2 |
| 2016 | Showing versus doing: Teaching by demonstrationabstractPeople often learn from others' demonstrations, and classic inverse reinforcement learning (IRL) algorithms have brought us closer to realizing this capacity in machines. In contrast, teaching by demonstration has been less well studied computationally. Here, we develop a novel Bayesian model for teaching by demonstration. Stark differences arise when demonstrators are intentionally teaching a task versus simply performing a task. In two experiments, we show that human participants systematically modify their teaching behavior consistent with the predictions of our model. Further, we show that even standard IRL algorithms benefit when learning from behaviors that are intentionally pedagogical. We conclude by discussing IRL algorithms that can take advantage of intentional pedagogy. Mark K. Ho, Michael L. Littman, James MacGlashan, Fiery Cushman, Joseph L. Austerweil |
NIPS | 1 |
| 2015 | Teaching with Rewards and Punishments: Reinforcement or Communication?
Mark K. Ho, Michael L. Littman, Fiery Cushman, Joseph L. Austerweil |
CogSci | 1 |
| 2013 | Working Memory and Abstract Representation in the Context of Culture
Mark K. Ho, Fiery Cushman |
CogSci | 1 |