VLDB 2026 Research / reviewers in the wild / expert
Serena Booth
dblp:195/6925
· DBLP profile ↗
13ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-7738-4418ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Should Robots Comply with Our Instructions or Intentions?abstractWhen people communicate, they often express their intent imperfectly, and human collaborators routinely compensate for these mistakes without issue. For example, if Alice asks for a spatula while serving soup, Bob may infer her intent and bring a ladle instead. This raises a key question for human–robot collaboration: should robots follow instructions literally or should they infer and act on human intent? We study how people expect robots to respond to ambiguous or incorrect instructions in collaborative kitchen scenarios. In this user study, participants either act directly on behalf of a robot or indirectly in observing a robot that may depart from literal instructions to pursue the inferred intent. We find that people generally prefer robots to take some action rather than refuse to comply, although people expect robots to attempt to satisfy the literal instruction (i.e., by thoroughly searching the scene) before taking an imperfect action to satisfy the intent. As large language models (LLMs) are increasingly used to model common sense, we conduct a pilot study to assess whether LLMs make the same decisions as human users about when robots should reinterpret requests. Tiffany Horter, Andrew Markham, Agathoniki Trigoni, Serena Booth |
HRI | 4 |
| 2026 | Human-Interactive Robot Learning: Definition, Challenges, and RecommendationsabstractRobot learning from humans has been proposed and researched for several decades as a means to enable robots to learn new skills or adapt existing ones to new situations. Recent advances in AI, including learning approaches like reinforcement learning and architectures like transformers and foundation models, combined with access to massive datasets, have created attractive opportunities to apply those data-hungry techniques to this problem. We argue that the focus on massive amounts of pre-collected data, and the resulting learning paradigm, where humans demonstrate and robots learn in isolation, is overshadowing a specialized area of work we term Human-Interactive Robot Learning (HIRL). This paradigm, wherein robots and humans interact during the learning process , is at the intersection of multiple fields (AI, robotics, human–computer interaction, design and others) and holds unique promise. Using HIRL, robots can achieve greater sample efficiency (as humans can provide task knowledge through interaction), align with human preferences (as humans can guide the robot behavior toward their expectations), and explore more meaningfully and safely (as humans can utilize domain knowledge to guide learning and prevent catastrophic failures). This can result in robotic systems that can more quickly and easily adapt to new tasks in human environments. The objective of this article is to provide a broad and consistent overview of HIRL research and to guide researchers toward understanding the scope of HIRL, and current open or underexplored challenges related to four themes—namely, human, robot learning, interaction, and broader context. The article includes concrete use cases to illustrate the interaction between these challenges and inspire further research according to broad recommendations and a call for action for the growing HIRL community. Kim Baraka, Ifrah Idrees, Taylor Kessler Faulkner, Erdem Biyik, Serena Booth, Mohamed Chetouani, Daniel H. Grollman, Akanksha Saran, Emmanuel Senft, Silvia Tulli, Anna-Lisa Vollmer, Antonio Andriella, Helen Beierling, Tiffany Horter, Jens Kober, Isaac S. Sheidlower, Matthew E. Taylor, Sanne van Waveren, Xuesu Xiao |
ACM Trans. Hum. Robot Interact. | 5 |
| 2025 | AI Governance and Lessons Learned as an AI Policy Advisor in the United States SenateabstractThis talk examines the intersection of artificial intelligence and policymaking, focusing on legislative and regulatory frameworks in the United States. It explores the role of key federal agencies, existing technology-agnostic laws affecting AI, and gaps in regulatory oversight that require legislative intervention. Consumer protection laws are analyzed for their relevance to AI governance, particularly in financial services. The discussion also highlights the implications for AI research, emphasizing the importance of interdisciplinary collaboration between computer scientists and policymakers to ensure responsible AI development that aligns with democratic values and societal interests. Serena Booth |
AAAI | 1 |
| 2025 | Goals vs. Rewards: Towards a Comparative Study of Objective Specification MechanismsabstractIn this late-breaking report, we look at two popular objective specification mechanisms for sequential decision-making problems, namely goals and rewards, and investigate how easy it would be for non-AI experts to use them effectively. Specifically, we propose a user study that allows us to test a user's ability to ($a$) use these mechanisms to direct a robot to generate some desired behavior and (b) predict the behavior resulting from a given objective specification. We conducted a small pilot study to test the study design and report some preliminary observations made regarding the two specification mechanisms. Septia Rani, Serena Booth, Sarath Sreedharan |
HRI | 2 |
| 2024 | Quality-Diversity Generative Sampling for Learning with Synthetic DataabstractGenerative models can serve as surrogates for some real data sources by creating synthetic training datasets, but in doing so they may transfer biases to downstream tasks. We focus on protecting quality and diversity when generating synthetic training datasets. We propose quality-diversity generative sampling (QDGS), a framework for sampling data uniformly across a user-defined measure space, despite the data coming from a biased generator. QDGS is a model-agnostic framework that uses prompt guidance to optimize a quality objective across measures of diversity for synthetically generated data, without fine-tuning the generative model. Using balanced synthetic datasets generated by QDGS, we first debias classifiers trained on color-biased shape datasets as a proof-of-concept. By applying QDGS to facial data synthesis, we prompt for desired semantic concepts, such as skin tone and age, to create an intersectional dataset with a combined blend of visual features. Leveraging this balanced data for training classifiers improves fairness while maintaining accuracy on facial recognition benchmarks. Code available at: https://github.com/Cylumn/qd-generative-sampling. Allen Chang, Matthew C. Fontaine, Serena Booth, Maja J. Mataric, Stefanos Nikolaidis |
AAAI | 3 |
| 2024 | Learning Optimal Advantage from Preferences and Mistaking It for RewardabstractWe consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most recent work assumes that human preferences are generated based only upon the reward accrued within those segments, or their partial return. Recent work casts doubt on the validity of this assumption, proposing an alternative preference model based upon regret. We investigate the consequences of assuming preferences are based upon partial return when they actually arise from regret. We argue that the learned function is an approximation of the optimal advantage function, not a reward function. We find that if a specific pitfall is addressed, this incorrect assumption is not particularly harmful, resulting in a highly shaped reward function. Nonetheless, this incorrect usage of the approximation of the optimal advantage function is less desirable than the appropriate and simpler approach of greedy maximization of it. From the perspective of the regret preference model, we also provide a clearer interpretation of fine tuning contemporary large language models with RLHF. This paper overall provides insight regarding why learning under the partial return preference model tends to work so well in practice, despite it conforming poorly to how humans give preferences. W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur O. Adalgeirsson, Serena Booth, Anca D. Dragan, Peter Stone 0001, Scott Niekum |
AAAI | 4 |
| 2023 | The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsabstractIn reinforcement learning (RL), a reward function that aligns exactly with a task's true performance metric is often necessarily sparse. For example, a true task metric might encode a reward of 1 upon success and 0 otherwise. The sparsity of these true task metrics can make them hard to learn from, so in practice they are often replaced with alternative dense reward functions. These dense reward functions are typically designed by experts through an ad hoc process of trial and error. In this process, experts manually search for a reward function that improves performance with respect to the task metric while also enabling an RL algorithm to learn faster. This process raises the question of whether the same reward function is optimal for all algorithms, i.e., whether the reward function can be overfit to a particular algorithm. In this paper, we study the consequences of this wide yet unexamined practice of trial-and-error reward design. We first conduct computational experiments that confirm that reward functions can be overfit to learning algorithms and their hyperparameters. We then conduct a controlled observation study which emulates expert practitioners' typical experiences of reward design, in which we similarly find evidence of reward function overfitting. We also find that experts' typical approach to reward design---of adopting a myopic strategy and weighing the relative goodness of each state-action pair---leads to misdesign through invalid task specifications, since RL algorithms use cumulative reward rather than rewards for individual state-action pairs as an optimization target. Code, data: github.com/serenabooth/reward-design-perils Serena Booth, W. Bradley Knox, Julie A. Shah, Scott Niekum, Peter Stone 0001, Alessandro Allievi |
AAAI | 1 |
| 2022 | Do Feature Attribution Methods Correctly Attribute Features?abstractFeature attribution methods are popular in interpretable machine learning. These methods compute the attribution of each input feature to represent its importance, but there is no consensus on the definition of "attribution", leading to many competing methods with little systematic evaluation, complicated in particular by the lack of ground truth attribution. To address this, we propose a dataset modification procedure to induce such ground truth. Using this procedure, we evaluate three common methods: saliency maps, rationales, and attentions. We identify several deficiencies and add new perspectives to the growing body of evidence questioning the correctness and reliability of these methods applied on datasets in the wild. We further discuss possible avenues for remedy and recommend new attribution methods to be tested against ground truth before deployment. The code and appendix are available at https://yilunzhou.github.io/feature-attribution-evaluation/. Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie A. Shah |
AAAI | 2 |
| 2022 | Revisiting Human-Robot Teaching and Learning Through the Lens of Human Concept LearningabstractWhen interacting with a robot, humans form con-ceptual models (of varying quality) which capture how the robot behaves. These conceptual models form just from watching or in-teracting with the robot, with or without conscious thought. Some methods select and present robot behaviors to improve human conceptual model formation; nonetheless, these methods and HRI more broadly have not yet consulted cognitive theories of human concept learning. These validated theories offer concrete design guidance to support humans in developing conceptual models more quickly, accurately, and flexibly. Specifically, Analogical Transfer Theory and the Variation Theory of Learning have been successfully deployed in other fields, and offer new insights for the HRI community about the selection and presentation of robot behaviors. Using these theories, we review and contextualize 35 prior works in human-robot teaching and learning, and we assess how these works incorporate or omit the design implications of these theories. From this review, we identify new opportunities for algorithms and interfaces to help humans more easily learn conceptual models of robot behaviors, which in turn can help humans become more effective robot teachers and collaborators. Serena Booth, Sanjana Sharma, Sarah Chung, Julie A. Shah, Elena L. Glassman |
HRI | 1 |
| 2021 | Bayes-TrEx: a Bayesian Sampling Approach to Model Transparency by ExampleabstractPost-hoc explanation methods are gaining popularity for interpreting, understanding, and debugging neural networks. Most analyses using such methods explain decisions in response to inputs drawn from the test set. However, the test set may have few examples that trigger some model behaviors, such as high-confidence failures or ambiguous classifications. To address these challenges, we introduce a flexible model inspection framework: Bayes-TrEx. Given a data distribution, Bayes-TrEx finds in-distribution examples which trigger a specified prediction confidence. We demonstrate several use cases of Bayes-TrEx, including revealing highly confident (mis)classifications, visualizing class boundaries via ambiguous examples, understanding novel-class extrapolation behavior, and exposing neural network overconfidence. We use Bayes-TrEx to study classifiers trained on CLEVR, MNIST, and Fashion-MNIST, and we show that this framework enables more flexible holistic model analysis than just inspecting the test set. Code and supplemental material are available at https://github.com/serenabooth/Bayes-TrEx. Serena Booth, Yilun Zhou, Ankit Shah 0003, Julie A. Shah |
AAAI | 1 |
| 2021 | Machine Learning Practices Outside Big Tech: How Resource Constraints Challenge Responsible DevelopmentabstractPractitioners from diverse occupations and backgrounds are increasingly using machine learning (ML) methods. Nonetheless, studies on ML Practitioners typically draw populations from Big Tech and academia, as researchers have easier access to these communities. Through this selection bias, past research often excludes the broader, lesser-resourced ML community---for example, practitioners working at startups, at non-tech companies, and in the public sector. These practitioners share many of the same ML development difficulties and ethical conundrums as their Big Tech counterparts; however, their experiences are subject to additional under-studied challenges stemming from deploying ML with limited resources, increased existential risk, and absent access to in-house research teams. We contribute a qualitative analysis of 17 interviews with stakeholders from organizations which are less represented in prior studies. We uncover a number of tensions which are introduced or exacerbated by these organizations' resource constraints---tensions between privacy and ubiquity, resource management and performance optimization, and access and monopolization. Increased academic focus on these practitioners can facilitate a more holistic understanding of ML limitations, and so is useful for prescribing a research agenda to facilitate responsible ML development for all. Aspen K. Hopkins, Serena Booth |
AIES | 2 |
| 2019 | Evaluating the Interpretability of the Knowledge Compilation Map: Communicating Logical Statements EffectivelyabstractKnowledge compilation techniques translate propositional theories into equivalent forms to increase their computational tractability. But, how should we best present these propositional theories to a human? We analyze the standard taxonomy of propositional theories for relative interpretability across three model domains: highway driving, emergency triage, and the chopsticks game. We generate decision-making agents which produce logical explanations for their actions and apply knowledge compilation to these explanations. Then, we evaluate how quickly, accurately, and confidently users comprehend the generated explanations. We find that domain, formula size, and negated logical connectives significantly affect comprehension while formula properties typically associated with interpretability are not strong predictors of human ability to comprehend the theory. Serena Booth, Christian J. Muise, Julie A. Shah |
IJCAI | 1 |
| 2017 | Piggybacking Robots: Human-Robot Overtrust in University Dormitory SecurityabstractCan overtrust in robots compromise physical security? We conducted a series of experiments in which a robot positioned outside a secure-access student dormitory asked passersby to assist it to gain access. We found individual participants were as likely to assist the robot in exiting the dormitory (40% assistance rate, 4/10 individuals) as in entering (19%, 3/16 individuals). Groups of people were more likely than individuals to assist the robot in entering (71%, 10/14 groups). When the robot was disguised as a food delivery agent for the fictional start-up Robot Grub, individuals were more likely to assist the robot in entering (76%, 16/21 individuals). Lastly, we found participants who identified the robot as a bomb threat demonstrated a trend toward assisting the robot (87%, 7/8 individuals, 6/7 groups). Thus, we demonstrate that overtrust---the unfounded belief that the robot does not intend to deceive or carry risk---can represent a significant threat to physical security at a university dormitory. Serena Booth, James Tompkin 0001, Hanspeter Pfister, Jim Waldo, Krzysztof Z. Gajos, Radhika Nagpal |
HRI | 1 |