VLDB 2026 Research / reviewers in the wild / expert
Jonathan Dodge
dblp:50/6860 · also Jonathan E. Dodge
· DBLP profile ↗
17ranked-venue papers
5as first author
6since 2021 · last 2026
0009-0000-5175-5962ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
3 papers |
Human-AI interaction · 83% Ubiquitous computing and smart environments · 9% Health and well-being technologies · 9% | |
| Artificial intelligence
3 papers |
Trustworthy machine learning · 68% Reinforcement learning · 22% Multi-agent systems · 10% |
Topics — the 5 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Human-AI interaction
explainable AI |
0.7 | 2 | 2019 | Explaining Reinforcement Learning to Mere Mortals: An Empirical Study · IJCAI 2019 How the Experts Do It: Assessing and Explaining Agent Behaviors in Real-Time Strategy Games · CHI 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2018 | Visualizing and Understanding Atari Agents · ICML 2018 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map |
0.3 | 1 | 2018 | Visualizing and Understanding Atari Agents · ICML 2018 |
Ubiquitous computing and smart environments › smart buildings
energy consumption feedback |
0.1 | 1 | 2010 | Studying always-on electricity feedback in the home · CHI 2010 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.1 | 1 | 2018 | Visualizing and Understanding Atari Agents · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
saliency map · 1.1reward decomposition · 0.8think-aloud protocol · 0.7qualitative analysis · 0.7contextual inquiry · 0.7user study · 0.1diary study · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI Can See What You Can't See: How LLM-Agents Complement Human-Based Gender-Inclusive Usability TestingabstractInclusive usability testing, such as the GenderMag method, wants to identify gender-related usability problems in digital interfaces. Large Language Models (LLMs) have been used by usability engineers in usability evaluations but their contribution is still underexplored, especially regarding inclusive usability testing. Research has shown that GenderMag workshops can produce valuable insights but are resource intense and might show effects from the evaluators’ ability to embody personas with a different cognitive style. Therefore, we need to assess if LLM-agent based testing can aid human-based evaluations. This study evaluates an LLM-agent system for GenderMag persona-based usability testing, and compares its performance to traditional human-led evaluations. The agent system integrates GenderMag persona facets into three LLM-agents, which analyze usability issues of four web interfaces, three generic and one intentionally flawed interface with gender-related usability issues. We quantitatively and qualitatively compare the types, severity, and relevance of usability problems LLM-agents identified to those produced in three GenderMag workshops involving nine participants. Findings show a broad overlap in detecting usability issues of humans and LLM-agents that are not gender specific via the generic interfaces. Agents thereby consistently assign significantly higher severity and relevance ratings. For the intentionally flawed interface, humans and LLM-agents assign similar ratings but the overlap between human and agent gender-related usability issues was low, with each missing issues the other caught. The agent system is an efficient tool for gender-related usability evaluation that humans may overlook, thereby expanding the coverage of evaluation. Joy Geuenich, Frank Joublin, Antonello Ceravola, Stefan Brandenburg, Jonathan Dodge |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2026 | Inclusive Design of AI's Explanations: Just for Those Previously Left Out?abstractAbstract Motivations . Explainable AI (XAI) systems aim to improve users’ understanding of AI, but XAI research has shown that many XAI explanations serve some users well while failing others. In non-AI systems, software practitioners have used inclusive design approaches to address similar problems, sometimes creating “curb-cut” improvements that benefit both underserved users and everyone else. This raises the possibility that inclusive design approaches can bring similar curb-cut improvements to AI explanations. Objectives . Our objective was to investigate possible curb-cut effects of inclusivity-driven fixes an AI product team made using an inclusive design approach (GenderMag) to improve their XAI prototype. Methods . We ran a between-subject study with 69 participants who had no formal AI background. 34 participants used the original version of the XAI prototype and the rest used the version with the AI team’s inclusivity fixes. We then compared the two groups’ mental model concepts scores and prediction accuracy, and the two prototypes’ inclusivity. Results . Our investigation produced four main results. First, the AI team’s inclusivity fixes were overall effective, resulting in overall better conceptual mental models with the new prototype. Further (second), the AI team’s inclusivity fixes were particularly beneficial to the underserved population’s conceptual mental models—which, together with the first result, constitutes a curb-cut effect. However (third), the inclusivity fixes did not improve participants’ prediction accuracy scores. Instead, it appears to have harmed them overall—a “curb-fence” effect (opposite of a curb-cut effect). Finally (fourth), the AI team’s fixes improved equity, reducing the gender gap by 45%. Md Montaser Hamid, Fatima A. Moussaoui, Jimena Noa Guevara, Andrew Anderson 0002, Puja Agarwal, Jonathan Dodge, Margaret M. Burnett |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2022 | How Do People Rank Multiple Mutant Agents?abstractFaced with several AI-powered sequential decision-making systems, how might someone choose on which to rely? For example, imagine car buyer Blair shopping for a self-driving car, or developer Dillon trying to choose an appropriate ML model to use in their application. Their first choice might be infeasible (i.e., too expensive in money or execution time), so they may need to select their second or third choice. To address this question, this paper presents: 1) Explanation Resolution, a quantifiable direct measurement concept; 2) a new XAI empirical task to measure explanations: “the Ranking Task”; and 3) a new strategy for inducing controllable agent variations—Mutant Agent Generation. In support of those main contributions, it also presents 4) novel explanations for sequential decision-making agents; 5) an adaptation to the AAR/AI assessment process; and 6) a qualitative study around these devices with 10 participants to investigate how they performed the Ranking Task on our mutant agents, using our explanations, and structured by AAR/AI. From an XAI researcher perspective, just as mutation testing can be applied to any code, mutant agent generation can be applied to essentially any neural network for which one wants to evaluate an assessment process or explanation type. As to an XAI user’s perspective, the participants ranked the agents well overall, but showed the importance of high explanation resolution for close differences between agents. The participants also revealed the importance of supporting a wide diversity of explanation diets and agent “test selection” strategies. Jonathan Dodge, Andrew Anderson 0002, Matthew L. Olson, Rupika Dikkala, Margaret M. Burnett |
IUI | 1 |
| 2022 | Finding AI's Faults with AAR/AI: An Empirical StudyabstractWould you allow an AI agent to make decisions on your behalf? If the answer is “not always,” the next question becomes “in what circumstances”? Answering this question requires human users to be able to assess an AI agent—and not just with overall pass/fail assessments or statistics. Here users need to be able to localize an agent’s bugs so that they can determine when they are willing to rely on the agent and when they are not. After-Action Review for AI (AAR/AI), a new AI assessment process for integration with Explainable AI systems, aims to support human users in this endeavor, and in this article we empirically investigate AAR/AI’s effectiveness with domain-knowledgeable users. Our results show that AAR/AI participants not only located significantly more bugs than non-AAR/AI participants did (i.e., showed greater recall) but also located them more precisely (i.e., with greater precision). In fact, AAR/AI participants outperformed non-AAR/AI participants on every bug and were, on average, almost six times as likely as non-AAR/AI participants to find any particular bug. Finally, evidence suggests that incorporating labeling into the AAR/AI process may encourage domain-knowledgeable users to abstract above individual instances of bugs; we hypothesize that doing so may have contributed further to AAR/AI participants’ effectiveness. Roli Khanna, Jonathan Dodge, Andrew Anderson 0002, Rupika Dikkala, Jed Irvine, Zeyad Shureih, Kin-Ho Lam, Caleb R. Matthews, Zhengxian Lin, Minsuk Kahng, Alan Fern, Margaret M. Burnett |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2021 | After-Action Review for AI (AAR/AI)abstractExplainable AI is growing in importance as AI pervades modern society, but few have studied how explainable AI can directly support people trying to assess an AI agent. Without a rigorous process, people may approach assessment in ad hoc ways—leading to the possibility of wide variations in assessment of the same agent due only to variations in their processes. AAR, or After-Action Review, is a method some military organizations use to assess human agents, and it has been validated in many domains. Drawing upon this strategy, we derived an After-Action Review for AI (AAR/AI), to organize ways people assess reinforcement learning agents in a sequential decision-making environment. We then investigated what AAR/AI brought to human assessors in two qualitative studies. The first investigated AAR/AI to gather formative information, and the second built upon the results, and also varied the type of explanation (model-free vs. model-based) used in the AAR/AI process. Among the results were the following: (1) participants reporting that AAR/AI helped to organize their thoughts and think logically about the agent, (2) AAR/AI encouraged participants to reason about the agent from a wide range of perspectives , and (3) participants were able to leverage AAR/AI with the model-based explanations to falsify the agent’s predictions. Jonathan Dodge, Roli Khanna, Jed Irvine, Kin-Ho Lam, Theresa Mai, Zhengxian Lin, Nicholas Kiddle, Evan Newman, Andrew Anderson 0002, Sai Raja, Caleb R. Matthews, Christopher Perdriau, Margaret M. Burnett, Alan Fern |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2021 | The Shoutcasters, the Game Enthusiasts, and the AI: Foraging for Explanations of Real-time Strategy PlayersabstractAssessing and understanding intelligent agents is a difficult task for users who lack an AI background. “Explainable AI” (XAI) aims to address this problem, but what should be in an explanation? One route toward answering this question is to turn to theories of how humans try to obtain information they seek. Information Foraging Theory (IFT) is one such theory. In this article, we present a series of studies 1 using IFT: the first investigates how expert explainers supply explanations in the RTS domain, the second investigates what explanations domain experts demand from agents in the RTS domain, and the last focuses on how both populations try to explain a state-of-the-art AI. Our results show that RTS environments like StarCraft offer so many options that change so rapidly, foraging tends to be very costly. Ways foragers attempted to manage such costs included “satisficing” approaches to reduce their cognitive load, such as focusing more on What information than on Why information, strategic use of language to communicate a lot of nuanced information in a few words, and optimizing their environment when possible to make their most valuable information patches readily available. Further, when a real AI entered the picture, even very experienced domain experts had difficulty understanding and judging some of the AI’s unconventional behaviors. Finally, our results reveal ways Information Foraging Theory can inform future XAI interactive explanation environments, and also how XAI can inform IFT. Sean Penney, Jonathan Dodge, Andrew Anderson 0002, Claudia Hilderbrand, Logan Simpson, Margaret M. Burnett |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2020 | Keeping it "organized and logical": after-action review for AI (AAR/AI)abstractExplainable AI (XAI) is growing in importance as AI pervades modern society, but few have studied how XAI can directly support people trying to assess an AI agent. Without a rigorous process, people may approach assessment in ad hoc ways---leading to the possibility of wide variations in assessment of the same agent due only to variations in their processes. AAR, or After-Action Review, is a method some military organizations use to assess human agents, and it has been validated in many domains. Drawing upon this strategy, we derived an AAR for AI, to organize ways people assess reinforcement learning (RL) agents in a sequential decision-making environment. The results of our qualitative study revealed several strengths and weaknesses of the AAR/AI process and the explanations embedded within it. Theresa Mai, Roli Khanna, Jonathan Dodge, Jed Irvine, Kin-Ho Lam, Zhengxian Lin, Nicholas Kiddle, Evan Newman, Sai Raja, Caleb R. Matthews, Christopher Perdriau, Margaret M. Burnett, Alan Fern |
IUI | 3 |
| 2020 | Mental Models of Mere Mortals with Explanations of Reinforcement LearningabstractHow should reinforcement learning (RL) agents explain themselves to humans not trained in AI? To gain insights into this question, we conducted a 124-participant, four-treatment experiment to compare participants’ mental models of an RL agent in the context of a simple Real-Time Strategy (RTS) game. The four treatments isolated two types of explanations vs. neither vs. both together. The two types of explanations were as follows: (1) saliency maps (an “Input Intelligibility Type” that explains the AI’s focus of attention) and (2) reward-decomposition bars (an “Output Intelligibility Type” that explains the AI’s predictions of future types of rewards). Our results show that a combined explanation that included saliency and reward bars was needed to achieve a statistically significant difference in participants’ mental model scores over the no-explanation treatment. However, this combined explanation was far from a panacea: It exacted disproportionately high cognitive loads from the participants who received the combined explanation. Further, in some situations, participants who saw both explanations predicted the agent’s next action worse than all other treatments’ participants. Andrew Anderson 0002, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Matthew L. Olson, Alan Fern, Margaret M. Burnett |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2019 | Explaining Reinforcement Learning to Mere Mortals: An Empirical StudyabstractWe present a user study to investigate the impact of explanations on non-experts? understanding of reinforcement learning (RL) agents. We investigate both a common RL visualization, saliency maps (the focus of attention), and a more recent explanation type, reward-decomposition bars (predictions of future types of rewards). We designed a 124 participant, four-treatment experiment to compare participants? mental models of an RL agent in a simple Real-Time Strategy (RTS) game. Our results show that the combination of both saliency and reward bars were needed to achieve a statistically significant improvement in mental model score over the control. In addition, our qualitative analysis of the data reveals a number of effects for further study. Andrew Anderson 0002, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Alan Fern, Margaret M. Burnett |
IJCAI | 2 |
| 2019 | Explaining models: an empirical study of how explanations impact fairness judgmentabstractEnsuring fairness of machine learning systems is a human-in-the-loop process. It relies on developers, users, and the general public to identify fairness problems and make improvements. To facilitate the process we need effective, unbiased, and user-friendly explanations that people can confidently rely on. Towards that end, we conducted an empirical study with four types of programmatically generated explanations to understand how they impact people's fairness judgments of ML systems. With an experiment involving more than 160 Mechanical Turk workers, we show that: 1) Certain explanations are considered inherently less fair, while others can enhance people's confidence in the fairness of the algorithm; 2) Different fairness problems-such as model-wide fairness issues versus case-specific fairness discrepancies-may be more effectively exposed through different styles of explanation; 3) Individual differences, including prior positions and judgment criteria of algorithmic fairness, impact how people react to different styles of explanation. We conclude with a discussion on providing personalized and adaptive explanations to support fairness judgments of ML systems. Jonathan Dodge, Qingzi Vera Liao, Rachel K. E. Bellamy, Casey Dugan |
IUI | 1 |
| 2018 | How the Experts Do It: Assessing and Explaining Agent Behaviors in Real-Time Strategy GamesabstractHow should an AI-based explanation system explain an agent's complex behavior to ordinary end users who have no background in AI? Answering this question is an active research area, for if an AI-based explanation system could effectively explain intelligent agents' behavior, it could enable the end users to understand, assess, and appropriately trust (or distrust) the agents attempting to help them. To provide insights into this question, we turned to human expert explainers in the real-time strategy domain --"shoutcasters"-- to understand (1) how they foraged in an evolving strategy game in real time, (2) how they assessed the players' behaviors, and (3) how they constructed pertinent and timely explanations out of their insights and delivered them to their audience. The results provided insights into shoutcasters' foraging strategies for gleaning information necessary to assess and explain the players; a characterization of the types of implicit questions shoutcasters answered; and implications for creating explanations by using the patterns and abstraction levels these human experts revealed. Jonathan Dodge, Sean Penney, Claudia Hilderbrand, Andrew Anderson 0002, Margaret M. Burnett |
CHI | 1 |
| 2018 | Visualizing and Understanding Atari AgentsabstractWhile deep reinforcement learning (deep RL) agents are effective at maximizing rewards, it is often unclear what strategies they use to do so. In this paper, we take a step toward explaining deep RL agents through a case study using Atari 2600 environments. In particular, we focus on using saliency maps to understand how an agent learns and executes a policy. We introduce a method for generating useful saliency maps and use it to show 1) what strong agents attend to, 2) whether agents are making decisions for the right or wrong reasons, and 3) how agents evolve during learning. We also test our method on non-expert human subjects and find that it improves their ability to reason about these agents. Overall, our results show that saliency information can provide significant insight into an RL agent’s decisions and learning behavior. Sam Greydanus, Anurag Koul, Jonathan Dodge, Alan Fern |
ICML | 3 |
| 2018 | Toward Foraging for Understanding of StarCraft Agents: An Empirical StudyabstractAssessing and understanding intelligent agents is a difficult task for users that lack an AI background. A relatively new area, called "Explainable AI," is emerging to help address this problem, but little is known about how users would forage through information an explanation system might offer. To inform the development of Explainable AI systems, we conducted a formative study -- using the lens of Information Foraging Theory -- into how experienced users foraged in the domain of StarCraft to assess an agent. Our results showed that participants faced difficult foraging problems. These foraging problems caused participants to entirely miss events that were important to them, reluctantly choose to ignore actions they did not want to ignore, and bear high cognitive, navigation, and information costs to access the information they needed. Sean Penney, Jonathan Dodge, Claudia Hilderbrand, Andrew Anderson 0002, Logan Simpson, Margaret M. Burnett |
IUI | 2 |
| 2010 | Studying always-on electricity feedback in the homeabstractThe recent emphasis on sustainability has made consumers more aware of their responsibility for saving resources, in particular, electricity. Consumers can better understand how to save electricity by gaining awareness of their consumption beyond the typical monthly bill. We conducted a study to understand consumers' awareness of energy consumption in the home and to determine their requirements for an interactive, always-on interface for exploring data to gain awareness of home energy consumption. In this paper, we describe a three-stage approach to supporting electricity conservation routines: raise awareness, inform complex changes, and maintain sustainable routines. We then present the findings from our study to support design implications for energy consumption feedback interfaces. Yann Riche, Jonathan Dodge, Ronald A. Metoyer |
CHI | 2 |
| 2010 | Explaining how to play real-time strategy games
Ronald A. Metoyer, Simone Stumpf, Christoph Neumann 0003, Jonathan Dodge, Jill Cao, Aaron Schnabel |
Knowl. Based Syst. | 4 |
| 2009 | Implications for an exercise prescription authoring notationabstractCommunicating dynamic motion content, such as exercise, with a static medium, such as paper, is difficult. The technology exists for presenting 3D animated exercise content to patients, however, the tools for allowing exercise domain experts to effectively author the content do not exist. We conducted two formative studies with exercise science domain experts to discover the requirements for an exercise prescription authoring notation. Based on our findings, we implemented a software prototype and performed a think-aloud study to understand its strengths and weaknesses. The results of our studies have implications for any software solution aimed at the authoring of physical activity content. Jonathan Dodge, Ronald A. Metoyer, Katherine B. Gunter |
VL/HCC | 1 |
| 2006 | System Management for Grid-Enabling a Vibroacoustic Analysis ApplicationabstractSystem management aspects are described for the process of grid-enabling a vibroacoustic analysis application using the Globus Toolkit 3.2.1. This is the first step in a project intended to grid-enable a suite of tools being developed as a service-oriented enterprise architecture for spacecraft telemetry analysis. Many of the applications in the suite are compute intensive and would benefit from significantly improved performance. In this paper we show the advantage of using Globus to grid-enable a single tool in a vibroacoustic analysis flow, with the result that using as few as eleven nodes, that tool's runtime improved by a factor of eight. While communication overhead does affect performance, these results also indicate that coordinated communication and execution scheduling as part of workflow management would be able to significantly improve overall efficiency. In the larger context, our experience also shows that the service-oriented architecture approach, using grid computing tools, can provide a more flexible system design, in addition to improved performance and increased utilization of resources. We also provide some lessons learned in using the Globus Toolkit Brian Bentow, Jonathan Dodge, Aaron Homer, Christopher D. Moore, Robert M. Keller, Matthew Presley, Jorge Seidel, Craig A. Lee, Joseph Betser |
NOMS | 2 |