EDBT 2026 Demo / reviewers in the wild / expert
Rachel K. E. Bellamy
dblp:33/3184
· DBLP profile ↗
43ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0002-9403-2913ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 28 · 4 first-authorSoftware engineering, systems software and programming languages · 9 · 2 first-authorArtificial intelligence and machine learning · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 53% Question answering and dialogue systems · 22% Information extraction and text analysis · 10% | |
| Human-computer interaction and pervasive computing
13 papers |
Human-AI interaction · 49% Usability and user experience research · 21% Collaborative and social computing · 19% | |
| Software engineering, system software, and programming languages
15 papers |
Empirical software engineering · 38% Software maintenance and evolution · 25% Requirements engineering and software design · 24% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 100% |
Topics — the 30 heaviest of 53, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability › explainable AI
explanation methods |
1.0 | 2 | 2022 | AI Explainability 360: Impact and Design · AAAI 2022 AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models · J. Mach. Learn. Res. 2020 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 2 | 2022 | AI Explainability 360: Impact and Design · AAAI 2022 AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models · J. Mach. Learn. Res. 2020 |
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation |
0.6 | 2 | 2022 | AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models · J. Mach. Learn. Res. 2020 AI Explainability 360: Impact and Design · AAAI 2022 |
Empirical software engineering
developer studies |
0.4 | 3 | 2013 | How Programmers Debug, Revisited: An Information Foraging Theory Perspective · IEEE Trans. Software Eng. 2013 An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 Moving into a new software project landscape · ICSE (1) 2010 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.4 | 1 | 2019 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.4 | 1 | 2019 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Natural language and speech › Information extraction and text analysis
text classification |
0.4 | 1 | 2019 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Machine learning › Learning paradigms
weakly supervised learning |
0.4 | 1 | 2019 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Empirical software engineering › user behavior analysis
information foraging theory |
0.4 | 3 | 2013 | The whats and hows of programmers' foraging diets · CHI 2013 Reactive information foraging: an empirical investigation of theory-based recommender systems for programmers · CHI 2012 How Programmers Debug, Revisited: An Information Foraging Theory Perspective · IEEE Trans. Software Eng. 2013 |
Natural language and speech › Question answering and dialogue systems › conversational agents
conversational assistant |
0.3 | 1 | 2018 | Water Advisor - A Data-Driven, Multi-Modal, Contextual Assistant to Help With Water Usage Decisions · AAAI 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
explicable planning |
0.3 | 1 | 2018 | Visualizations for an Explainable Planning Agent · IJCAI 2018 |
Visualization and visual analytics › explainable AI › explainable machine learning
explanation visualization |
0.3 | 1 | 2018 | Visualizations for an Explainable Planning Agent · IJCAI 2018 |
Human-AI interaction
conversational agents |
0.3 | 1 | 2018 | Face Value? · CHI 2018 |
Human-AI interaction › conversational agents
embodied conversational agents |
0.3 | 1 | 2018 | Face Value? · CHI 2018 |
Collaborative and social computing › cooperative work
group decision-making |
0.3 | 1 | 2018 | Face Value? · CHI 2018 |
Human-AI interaction › human-in-the-loop
human-in-the-loop decision making |
0.3 | 1 | 2018 | Visualizations for an Explainable Planning Agent · IJCAI 2018 |
Requirements engineering and software design
model-driven engineering |
0.2 | 2 | 2011 | Workshop on flexible modeling tools: (FlexiTools 2011) · ICSE 2011 Flexible Modeling Tools (FlexiTools2010) · ICSE (2) 2010 |
Requirements engineering and software design
software architecture |
0.2 | 2 | 2011 | Workshop on flexible modeling tools: (FlexiTools 2011) · ICSE 2011 Flexible Modeling Tools (FlexiTools2010) · ICSE (2) 2010 |
Usability and user experience research
decision-making |
0.2 | 1 | 2015 | Designing Information for Remediating Cognitive Biases in Decision-Making · CHI 2015 |
Human-AI interaction
decision support |
0.2 | 1 | 2015 | Designing Information for Remediating Cognitive Biases in Decision-Making · CHI 2015 |
Collaborative and social computing › information seeking
information foraging |
0.2 | 2 | 2010 | Reactive information foraging for evolving goals · CHI 2010 Using information scent to model the dynamic foraging behavior of programmers in maintenance tasks · CHI 2008 |
Software maintenance and evolution › program comprehension
code navigation |
0.2 | 2 | 2010 | Reactive information foraging for evolving goals · CHI 2010 Using information scent to model the dynamic foraging behavior of programmers in maintenance tasks · CHI 2008 |
Software maintenance and evolution
refactoring |
0.2 | 1 | 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 |
Software maintenance and evolution › refactoring
refactoring tool support |
0.2 | 1 | 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 |
Software testing
regression testing |
0.2 | 1 | 2013 | Human performance regression testing · ICSE 2013 |
Usability and user experience research › human performance modeling
predictive human performance modeling |
0.1 | 1 | 2012 | Easing the generation of predictive human performance models from legacy systems · CHI 2012 |
User interface design and tools › design tools
sketching tools |
0.1 | 1 | 2011 | Sketching tools for ideation · ICSE 2011 |
Natural language and speech › Information extraction and text analysis › data annotation
data programming |
0.1 | 1 | 2019 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Empirical software engineering › developer studies › software teams
newcomer onboarding |
0.1 | 1 | 2010 | Moving into a new software project landscape · ICSE (1) 2010 |
Usability and user experience research › usability evaluation
usability testing |
0.0 | 1 | 2013 | Human performance regression testing · ICSE 2013 |
Methods — techniques the papers use, named apart from their topics
multimodal AI · 1.0automated planning · 1.0information foraging theory · 0.9empirical study · 0.6explainability methods · 0.6qualitative analysis · 0.5expected return method · 0.4bootstrapping · 0.4taxonomy · 0.4software toolkit · 0.4controlled experiment · 0.4weak supervision · 0.4search-label-propagate · 0.4data programming · 0.4inference algorithm · 0.3cognitive modeling · 0.3sampling · 0.2test case generation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | AI Explainability 360: Impact and DesignabstractAs artificial intelligence and machine learning algorithms become increasingly prevalent in society, multiple stakeholders are calling for these algorithms to provide explanations. At the same time, these stakeholders, whether they be affected citizens, government regulators, domain experts, or system developers, have different explanation needs. To address these needs, in 2019, we created AI Explainability 360, an open source software toolkit featuring ten diverse and state-of-the-art explainability methods and two evaluation metrics. This paper examines the impact of the toolkit with several case studies, statistics, and community feedback. The different ways in which users have experienced AI Explainability 360 have resulted in multiple types of impact and improvements in multiple metrics, highlighted by the adoption of the toolkit by the independent LF AI & Data Foundation. The paper also describes the flexible design of the toolkit, examples of its use, and the significant educational material and documentation available to its users. Vijay Arya, Rachel K. E. Bellamy, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Qingzi Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam 0001, Moninder Singh, Kush R. Varshney, Dennis Wei |
AAAI | 2 |
| 2020 | Joint Optimization of AI Fairness and Utility: A Human-Centered ApproachabstractToday, AI is increasingly being used in many high-stakes decision-making applications in which fairness is an important concern. Already, there are many examples of AI being biased and making questionable and unfair decisions. The AI research community has proposed many methods to measure and mitigate unwanted biases, but few of them involve inputs from human policy makers. We argue that because different fairness criteria sometimes cannot be simultaneously satisfied, and because achieving fairness often requires sacrificing other objectives such as model accuracy, it is key to acquire and adhere to human policy makers' preferences on how to make the tradeoff among these objectives. In this paper, we propose a framework and some exemplar methods for eliciting such preferences and for optimizing an AI model according to these preferences. Rachel K. E. Bellamy, Kush R. Varshney |
AIES | 2 |
| 2020 | Towards Designing Conversational Agents for Pair Programming: Accounting for Creativity Strategies and Conversational StylesabstractEstablished research on pair programming reveals benefits, including increasing communication, creativity, self-efficacy, and promoting gender inclusivity. However, research has reported limitations such as finding a compatible partner, scheduling sessions between partners, and resistance to pairing. Further, pairings can be affected by predispositions to negative stereotypes. These problems can be addressed by replacing one human member of the pair with a conversational agent. To investigate the design space of such a conversational agent, we conducted a controlled remote pair programming study. Our analysis found various creative problem-solving strategies and differences in conversational styles. We further analyzed the transferable strategies from human-human collaboration to human-agent collaboration by conducting a Wizard of Oz study. The findings from the two studies helped us gain insights regarding design of a programmer conversational agent. We make recommendations for researchers and practitioners for designing pair programming conversational agent tools. Sandeep Kaur Kuttal, Jarow Myers, Sam Gurka, David Magar, David Piorkowski, Rachel K. E. Bellamy |
VL/HCC | 6 |
| 2020 | Can Machine Learning Facilitate Remote Pair Programming? Challenges, Insights & ImplicationsabstractRemote pair programming encapsulates the benefits of well-researched (co-located) pair programming. However, its effectiveness is hindered by challenges including pair incompatibility, imbalanced roles, and inclinations to work alone. Recent research has explored pedagogical methods to alleviate these challenges, but none have considered the integration of machine learning agents to facilitate remote pair programming. Therefore, we investigated the capabilities of popular text classification algorithms on identifying three facets of pair programming: dialogue acts, creativity stages, and pair programming roles. We collected a dataset of 3,436 utterances from a lab study of 18 pair programmers in a simulated remote environment. We found that pair programming dialogue poses a challenge as it is often unpremeditated and inadequately structured. Despite this, the accuracy of our machine learning classifier was improved by the choice of contextual dialogue features. Our results have implications for facilitating pair programming in global software development and online computer science education. Peter Robe, Sandeep Kaur Kuttal, Rachel K. E. Bellamy |
VL/HCC | 4 |
| 2020 | AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning ModelsabstractAs artificial intelligence algorithms make further inroads in high-stakes societal applications, there are increasing calls from multiple stakeholders for these algorithms to explain their outputs. To make matters more challenging, different personas of consumers of explanations have different requirements for explanations. Toward addressing these needs, we introduce AI Explainability 360, an open-source Python toolkit featuring ten diverse and state-of-the-art explainability methods and two evaluation metrics. Equally important, we provide a taxonomy to help entities requiring explanations to navigate the space of interpretation and explanation methods, not only those in the toolkit but also in the broader literature on explainability. For data scientists and other users of the toolkit, we have implemented an extensible software architecture that organizes methods according to their place in the AI modeling pipeline. The toolkit is not only the software, but also guidance material, tutorials, and an interactive web demo to introduce AI explainability to different audiences. Together, our toolkit and taxonomy can help identify gaps where more explainability methods are needed and provide a platform to incorporate them as they are developed. Vijay Arya, Rachel K. E. Bellamy, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Qingzi Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam 0001, Moninder Singh, Kush R. Varshney, Dennis Wei |
J. Mach. Learn. Res. | 2 |
| 2020 | Explainable Active Learning (XAL): Toward AI Explanations as Interfaces for Machine TeachersabstractThe wide adoption of Machine Learning (ML) technologies has created a growing demand for people who can train ML models. Some advocated the term "machine teacher'' to refer to the role of people who inject domain knowledge into ML models. This "teaching'' perspective emphasizes supporting the productivity and mental wellbeing of machine teachers through efficient learning algorithms and thoughtful design of human-AI interfaces. One promising learning paradigm is Active Learning (AL), by which the model intelligently selects instances to query a machine teacher for labels, so that the labeling workload could be largely reduced. However, in current AL settings, the human-AI interface remains minimal and opaque. A dearth of empirical studies further hinders us from developing teacher-friendly interfaces for AL algorithms. In this work, we begin considering AI explanations as a core element of the human-AI interface for teaching machines. When a human student learns, it is a common pattern to present one's own reasoning and solicit feedback from the teacher. When a ML model learns and still makes mistakes, the teacher ought to be able to understand the reasoning underlying its mistakes. When the model matures, the teacher should be able to recognize its progress in order to trust and feel confident about their teaching outcome. Toward this vision, we propose a novel paradigm of explainable active learning (XAL), by introducing techniques from the surging field of explainable AI (XAI) into an AL setting. We conducted an empirical study comparing the model learning outcomes, feedback content and experience with XAL, to that of traditional AL and coactive learning (providing the model's prediction without explanation). Our study shows benefits of AI explanation as interfaces for machine teaching--supporting trust calibration and enabling rich forms of teaching feedback, and potential drawbacks--anchoring effect with the model judgment and additional cognitive workload. Our study also reveals important individual factors that mediate a machine teacher's reception to AI explanations, including task knowledge, AI experience and Need for Cognition. By reflecting on the results, we suggest future directions and design implications for XAL, and more broadly, machine teaching through AI explanations. Bhavya Ghai, Qingzi Vera Liao, Rachel K. E. Bellamy, Klaus Mueller 0001 |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2019 | Exploring Rural Community Practices in HIV Management for the Design of Technology for Hypertensive Patients Living with HIVabstractInformation communication technologies for development (ICTD) can support people with chronic illnesses living in rural communities. In Kenya, ICTD use in areas where undetected cases of hypertension and high HIV infection rates exist is underexplored. Partnering with a health facility in Migori, Kenya, we report on the uses of technology in managing HIV. We see the use of technology to manage HIV was influenced by the roles and routines of patients and clinicians, trust between practitioners and patients, and sources of data that clinicians use for patient examination. We use these results to inform the design of technologies that can support patients living with comorbid HIV and hypertension, as well as their care providers, to manage their care in similar settings. We also reiterate the important mediatory role that community health volunteers (CHVs) can play in the adoption of technology as patients manage their condition(s) once out of hospital. Erick Oduor, Carolyn Pang, Charles Wachira, Rachel K. E. Bellamy, Timothy Nyota, Sekou L. Remy, Aisha Walcott-Bryant, Wycliffe Omwanda, Julius Mbeya |
Conference on Designing Interactive Systems | 4 |
| 2019 | Bootstrapping Conversational Agents with Weak SupervisionabstractMany conversational agents in the market today follow a standard bot development framework which requires training intent classifiers to recognize user input. The need to create a proper set of training examples is often the bottleneck in the development process. In many occasions agent developers have access to historical chat logs that can provide a good quantity as well as coverage of training examples. However, the cost of labeling them with tens to hundreds of intents often prohibits taking full advantage of these chat logs. In this paper, we present a framework called search, label, and propagate (SLP) for bootstrapping intents from existing chat logs using weak supervision. The framework reduces hours to days of labeling effort down to minutes of work by using a search engine to find examples, then relies on a data programming approach to automatically expand the labels. We report on a user study that shows positive user feedback for this new approach to build conversational agents, and demonstrates the effectiveness of using data programming for autolabeling. While the system is developed for training conversational agents, the framework has broader application in significantly reducing labeling effort for training text classifiers. Neil Mallinar, Abhishek Shah, Rajendra Ugrani, Manikandan Gurusankar, Tin Kam Ho, Qingzi Vera Liao, Rachel K. E. Bellamy, Robert Yates, Chris Desmarais, Blake McGregor |
AAAI | 9 |
| 2019 | Explaining models: an empirical study of how explanations impact fairness judgmentabstractEnsuring fairness of machine learning systems is a human-in-the-loop process. It relies on developers, users, and the general public to identify fairness problems and make improvements. To facilitate the process we need effective, unbiased, and user-friendly explanations that people can confidently rely on. Towards that end, we conducted an empirical study with four types of programmatically generated explanations to understand how they impact people's fairness judgments of ML systems. With an experiment involving more than 160 Mechanical Turk workers, we show that: 1) Certain explanations are considered inherently less fair, while others can enhance people's confidence in the fairness of the algorithm; 2) Different fairness problems-such as model-wide fairness issues versus case-specific fairness discrepancies-may be more effectively exposed through different styles of explanation; 3) Individual differences, including prior positions and judgment criteria of algorithmic fairness, impact how people react to different styles of explanation. We conclude with a discussion on providing personalized and adaptive explanations to support fairness judgments of ML systems. Jonathan Dodge, Qingzi Vera Liao, Rachel K. E. Bellamy, Casey Dugan |
IUI | 4 |
| 2018 | Water Advisor - A Data-Driven, Multi-Modal, Contextual Assistant to Help With Water Usage DecisionsabstractWe demonstrate Water Advisor, a multi-modal assistant to help non-experts make sense of complex water quality data and apply it to their specific needs. A user can chat with the tool about water quality and activities of interest, and the system tries to advise using available water data for a location, applicable water regulations and relevant parameters using AI methods. Jason B. Ellis, Biplav Srivastava, Rachel K. E. Bellamy, Andy Aaron |
AAAI | 3 |
| 2018 | Face Value?abstractWe are interested in increasing the ability of groups to collaborate efficiently by leveraging new advances in AI and Conversational Agent (CA) technology. Given the longstanding debate on the necessity of embodiment for CAs, bringing them to groups requires answering the questions of whether and how providing a CA with a face affects its interaction with the humans in a group. We explored these questions by comparing group decision-making sessions facilitated by an embodied agent, versus a voice-only agent. Results of an experiment with 20 user groups revealed that while the embodiment improved various aspects of group's social perception of the agent (e.g., rapport, trust, intelligence, and power), its impact on the group-decision process and outcome was nuanced. Drawing on both quantitative and qualitative findings, we discuss the pros and cons of embodiment, argue that the value of having a face depends on the types of assistance the agent provides, and lay out directions for future research. Ameneh Shamekhi, Qingzi Vera Liao, Dakuo Wang, Rachel K. E. Bellamy, Thomas Erickson |
CHI | 4 |
| 2018 | Visualizations for an Explainable Planning AgentabstractIn this demonstration, we report on the visualization capabilities of an Explainable AI Planning (XAIP) agent that can support human-in-the-loop decision-making. Imposing transparency and explainability requirements on such agents is crucial for establishing human trust and common ground with an end-to-end automated planning system. Visualizing the agent's internal decision making processes is a crucial step towards achieving this. This may include externalizing the "brain" of the agent: starting from its sensory inputs, to progressively higher order decisions made by it in order to drive its planning components. We demonstrate these functionalities in the context of a smart assistant in the Cognitive Environments Laboratory at IBM's T.J. Watson Research Center. Tathagata Chakraborti, Kshitij Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O. Kephart, Rachel K. E. Bellamy |
IJCAI | 7 |
| 2016 | Diagnostic visualization for non-expert machine learning practitioners: A design studyabstractAs machine learning (ML) becomes increasingly popular, developers without deep experience in ML - who we will refer to as ML practitioners - are facing the need to diagnose problems with ML models. Yet successful diagnosis requires high-level expertise that practitioners lack. As in many complex data-oriented domains, visualization could help. This two-phase study explored the design of visualizations to aid ML diagnosis. In phase 1, twelve ML practitioners were asked to diagnose a model using ten state-of-the-art visualizations; seven design themes were identified. In phase 2, several design themes were embodied in an interactive visualization. The visualization was used to engage practitioners in a participatory design exercise that explored how they would carry out multi-step diagnosis using the visualization. Our findings provide design implications for tools that better support ML diagnosis by non-expert practitioners. Rachel K. E. Bellamy, Peter K. Malkin, Thomas Erickson |
VL/HCC | 2 |
| 2016 | Trials and tribulations of developers of intelligent systems: A field studyabstractIntelligent systems are gaining in popularity and receiving increased media attention, but little is known about how people actually go about developing them. In this paper, we attempt to fill this gap through a set of field interviews that investigate how people develop intelligent systems that incorporate machine learning algorithms. The developers we interviewed were experienced at working with machine learning algorithms and dealing with the large amounts of data needed to develop intelligent systems. Despite their level of experience, we learned that they struggle to establish a repeatable process. They described problems with each step of the processes they perform, as well as cross-cutting issues that pervade multiple steps of their processes. The unique difficulties that developers like these face seem to point to a need for software engineering advances that address such machine learning systems, and we conclude by discussing this need and some of its implications. Charles Hill 0001, Rachel K. E. Bellamy, Thomas Erickson, Margaret M. Burnett |
VL/HCC | 2 |
| 2015 | Designing Information for Remediating Cognitive Biases in Decision-MakingabstractSoftware is playing an increasingly important role in supporting human decision-making. Previous HCI research on decision support systems (DSS) has improved the information visualization aspect of DSS information design, but has somewhat overlooked the cognitive aspect of decision-making, namely that human reasoning is heuristic and reflects systematic errors or cognitive biases. We report on an empirical study of two cognitive biases: conservatism and loss aversion. Two remediation techniques recommended by previous research were tested: the expected return method, an actuarial-inspired approach presenting objective metrics; and bootstrapping, a technique successful in improving judgment consistency. The results show that the two biases can occur simultaneously and can have a huge impact on decision-making. The results also show that the two debiasing techniques are only partly effective. These findings suggest a need for more research on debiasing, and indicate some directions for exploring debiasing techniques and building decision support systems. Rachel K. E. Bellamy, Wendy A. Kellogg |
CHI | 2 |
| 2013 | The whats and hows of programmers' foraging dietsabstractOne of the least studied areas of Information Foraging Theory is diet: the information foragers choose to seek. For example, do foragers choose solely based on cost, or do they stubbornly pursue certain diets regardless of cost? Do their debugging strategies vary with their diets? To investigate "what" and "how" questions like these for the domain of software debugging, we qualitatively analyzed 9 professional developers' foraging goals, goal patterns, and strategies. Participants spent 50% of their time foraging. Of their foraging, 58% fell into distinct dietary patterns - mostly in patterns not previously discussed in the literature. In general, programmers' foraging strategies leaned more heavily toward enrichment than we expected, but different strategies aligned with different goal types. These and our other findings help fill the gap as to what programmers' dietary goals are and how their strategies relate to those goals. David Piorkowski, Scott D. Fleming, Irwin Kwan, Margaret M. Burnett, Christopher Scaffidi, Rachel K. E. Bellamy, Joshua Jordahl |
CHI | 6 |
| 2013 | Human performance regression testingabstractAs software systems evolve, new interface features such as keyboard shortcuts and toolbars are introduced. While it is common to regression test the new features for functional correctness, there has been less focus on systematic regression testing for usability, due to the effort and time involved in human studies. Cognitive modeling tools such as CogTool provide some help by computing predictions of user performance, but they still require manual effort to describe the user interface and tasks, limiting regression testing efforts. In recent work, we developed CogTool-Helper to reduce the effort required to generate human performance models of existing systems. We build on this work by providing task specific test case generation and present our vision for human performance regression testing (HPRT) that generates large numbers of test cases and evaluates a range of human performance predictions for the same task. We examine the feasibility of HPRT on four tasks in LibreOffice, find several regressions, and then discuss how a project team could use this information. We also illustrate that we can increase efficiency with sampling by leveraging an inference algorithm. Samples that take approximately 50% of the runtime lose at most 10% of the performance predictions. Amanda Swearngin, Myra B. Cohen, Bonnie E. John, Rachel K. E. Bellamy |
ICSE | 4 |
| 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse TasksabstractTheories of human behavior are an important but largely untapped resource for software engineering research. They facilitate understanding of human developers’ needs and activities, and thus can serve as a valuable resource to researchers designing software engineering tools. Furthermore, theories abstract beyond specific methods and tools to fundamental principles that can be applied to new situations. Toward filling this gap, we investigate the applicability and utility of Information Foraging Theory (IFT) for understanding information-intensive software engineering tasks, drawing upon literature in three areas: debugging, refactoring, and reuse. In particular, we focus on software engineering tools that aim to support information-intensive activities, that is, activities in which developers spend time seeking information. Regarding applicability, we consider whether and how the mathematical equations within IFT can be used to explain why certain existing tools have proven empirically successful at helping software engineers. Regarding utility, we applied an IFT perspective to identify recurring design patterns in these successful tools, and consider what opportunities for future research are revealed by our IFT perspective. Scott D. Fleming, Christopher Scaffidi, David Piorkowski, Margaret M. Burnett, Rachel K. E. Bellamy, Joseph Lawrance, Irwin Kwan |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2013 | How Programmers Debug, Revisited: An Information Foraging Theory PerspectiveabstractMany theories of human debugging rely on complex mental constructs that offer little practical advice to builders of software engineering tools. Although hypotheses are important in debugging, a theory of navigation adds more practical value to our understanding of how programmers debug. Therefore, in this paper, we reconsider how people go about debugging in large collections of source code using a modern programming environment. We present an information foraging theory of debugging that treats programmer navigation during debugging as being analogous to a predator following scent to find prey in the wild. The theory proposes that constructs of scent and topology provide enough information to describe and predict programmer navigation during debugging, without reference to mental states such as hypotheses. We investigate the scope of our theory through an empirical study of 10 professional programmers debugging a real-world open source program. We found that the programmers' verbalizations far more often concerned scent-following than hypotheses. To evaluate the predictiveness of our theory, we created an executable model that predicted programmer navigation behavior more accurately than comparable models that did not consider information scent. Finally, we discuss the implications of our results for enhancing software engineering tools. Joseph Lawrance, Christopher Bogart, Margaret M. Burnett, Rachel K. E. Bellamy, Kyle Rector, Scott D. Fleming |
IEEE Trans. Software Eng. | 4 |
| 2012 | Reactive information foraging: an empirical investigation of theory-based recommender systems for programmersabstractInformation Foraging Theory (IFT) has established itself as an important theory to explain how people seek information, but most work has focused more on the theory itself than on how best to apply it. In this paper, we investigate how to apply a reactive variant of IFT (Reactive IFT) to design IFT-based tools, with a special focus on such tools for ill-structured problems. Toward this end, we designed and implemented a variety of recommender algorithms to empirically investigate how to help people with the ill-structured problem of finding where to look for information while debugging source code. We varied the algorithms based on scent type supported (words alone vs. words + code structure), and based on use of foraging momentum to estimate rapidity of foragers' goal changes. Our empirical results showed that (1) using both words and code structure significantly improved the ability of the algorithms to recommend where software developers should look for information; (2) participants used recommendations to discover new places in the code and also as shortcuts to navigate to known places; and (3) low-momentum recommendations were significantly more useful than high-momentum recommendations, suggesting rapid and numerous goal changes in this type of setting. Overall, our contributions include two new recommendation algorithms, empirical evidence about when and why participants found IFT-based recommendations useful, and implications for the design of tools based on Reactive IFT. David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Bonnie E. John, Rachel K. E. Bellamy, Calvin Swart |
CHI | 7 |
| 2012 | Easing the generation of predictive human performance models from legacy systemsabstractWith the rise of tools for predictive human performance modeling in HCI comes a need to model legacy applications. Models of legacy systems are used to compare products to competitors, or new proposed design ideas to the existing version of an application. We present CogTool-Helper, an exemplar of a tool that results from joining this HCI need to research in automatic GUI testing from the Software Engineering testing community. CogTool-Helper uses automatic UI-model extraction and test case generation to automatically create CogTool storyboards and models and infer methods to accomplish tasks beyond what the UI designer has specified. A design walkthrough with experienced CogTool users reveal that CogTool-Helper resonates with a "pain point" of real-world modeling and provide suggestions for future work. Amanda Swearngin, Myra B. Cohen, Bonnie E. John, Rachel K. E. Bellamy |
CHI | 4 |
| 2012 | Using the "Physics" of notations to analyze a visual representation of business decision modelingabstractVisual representations are common in communicating about artifacts such as computer programs, software architectures, and business rules. Yet, generally speaking, these representations seem much harder to learn and to use than many of the representations in other domains. Moody [1] has pointed this out and proposed a set of principles for visual representations based on a wide review of relevant literature in cognitive psychology and software engineering. The real test of this framework is to use it. In this paper, we apply the principles set forth in Moody to examine and improve a proposed representation for business rules and business decisions. John C. Thomas, Judah Diament, Jacquelyn Martino, Rachel K. E. Bellamy |
VL/HCC | 4 |
| 2011 | Sketching tools for ideationabstractSketching facilitates design in the exploration of ideas about concrete objects and abstractions. In fact, throughout the software engineering process when grappling with new ideas, people reach for a pen and start sketching. While pen and paper work well, digital media can provide additional features to benefit the sketcher. Digital support will only be successful, however, if it does not detract from the core sketching experience. Based on research that defines characteristics of sketches and sketching, this paper offers three preliminary tool examples. Each example is intended to enable sketching while maintaining its characteristic experience. Rachel K. E. Bellamy, Michael Desmond, Jacquelyn Martino, Paul Matchen, Harold Ossher, John T. Richards, Calvin Swart |
ICSE | 1 |
| 2011 | Deploying CogTool: integrating quantitative usability assessment into real-world software developmentabstractUsability concerns are often difficult to integrate into real-world software development processes. To remedy this situation, IBM research and development, partnering with Carnegie Mellon University, has begun to employ a repeatable and quantifiable usability analysis method, embodied in CogTool, in its development practice. CogTool analyzes tasks performed on an interactive system from a storyboard and a demonstration of tasks on that storyboard, and predicts the time a skilled user will take to perform those tasks. We discuss how IBM designers and UX professionals used CogTool in their existing practice for contract compliance, communication within a product team and between a product team and its customer, assigning appropriate personnel to fix customer complaints, and quantitatively assessing design ideas before a line of code is written. We then reflect on the lessons learned by both the development organizations and the researchers attempting this technology transfer from academic research to integration into real-world practice, and we point to future research to even better serve the needs of practice. Rachel K. E. Bellamy, Bonnie E. John, Sandra Kogan |
ICSE | 1 |
| 2011 | Workshop on flexible modeling tools: (FlexiTools 2011)abstractModeling tools are often not used for tasks during the software lifecycle for which they should be more helpful; instead free-from approaches, such as office tools and white boards, are frequently used. Prior workshops explored why this is the case and what might be done about it. The goal of this workshop is to continue those discussions and also to form an initial set of challenge problems and research challenges that researchers and developers of flexible modeling tools should address. Harold Ossher, André van der Hoek, Margaret-Anne D. Storey, John C. Grundy, Rachel K. E. Bellamy, Marian Petre |
ICSE | 5 |
| 2011 | Modeling programmer navigation: A head-to-head empirical evaluation of predictive modelsabstractSoftware developers frequently need to perform code maintenance tasks, but doing so requires time-consuming navigation through code. A variety of tools are aimed at easing this navigation by using models to identify places in the code that a developer might want to visit, and then providing shortcuts so that the developer can quickly navigate to those locations. To date, however, only a few of these models have been compared head-to-head to assess their predictive accuracy. In particular, we do not know which models are most accurate overall, which are accurate only in certain circumstances, and whether combining models could enhance accuracy. Therefore, we have conducted an empirical study to evaluate the accuracy of a broad range of models for predicting many different kinds of code navigations in sample maintenance tasks. Overall, we found that models tended to perform best if they took into account how recently a developer has viewed pieces of the code, and if models took into account the spatial proximity of methods within the code. We also found that the accuracy of single-factor models can be improved by combining factors, using a spreading-activation based approach, to produce multi-factor models. Based on these results, we offer concrete guidance about how these models could be used to provide enhanced software development tools that ease the difficulty of navigating through code. David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Liza John, Christopher Bogart, Bonnie E. John, Margaret M. Burnett, Rachel K. E. Bellamy |
VL/HCC | 8 |
| 2010 | Towards a tool for keystroke level modeling of skilled screen readingabstractDesigners often have no access to individuals who use screen reading software, and may have little understanding of how their design choices impact these users. We explore here whether cog-nitive models of auditory interaction could provide insight into screen reader usability. By comparing human data with a tool-generated model of a practiced task performed using a screen reader, we identify several requirements for such models and tools. Most important is the need to represent parallel execution of hearing with thinking and acting. Rules for placement of cogni-tive operators that were developed for visual user interfaces may not be applicable in the auditory domain. Other mismatches be-tween the data and the model were attributed to the extremely fast listening rate and differences between the typing patterns of screen reader usage and the model's assumptions. This work in-forms the development of more accurate models of auditory inter-action. Tools incorporating such models could help designers create user interfaces that are well tuned for screen reader users, without the need for modeling expertise. Shari Trewin, Bonnie E. John, John T. Richards, Calvin Swart, Jonathan P. Brezin, Rachel K. E. Bellamy, John C. Thomas |
ASSETS | 6 |
| 2010 | Reactive information foraging for evolving goalsabstractInformation foraging models have predicted the navigation paths of people browsing the web and (more recently) of programmers while debugging, but these models do not explicitly model users' goals evolving over time. We present a new information foraging model called PFIS2 that does model information seeking with potentially evolving goals. We then evaluated variants of this model in a field study that analyzed programmers' daily navigations over a seven-month period. Our results were that PFIS2 predicted users' navigation remarkably well, even though the goals of navigation, and even the information landscape itself, were changing markedly during the pursuit of information. Joseph Lawrance, Margaret M. Burnett, Rachel K. E. Bellamy, Christopher Bogart, Calvin Swart |
CHI | 3 |
| 2010 | Moving into a new software project landscapeabstractWhen developers join a software development project, they find themselves in a project landscape, and they must become familiar with the various landscape features. To better understand the nature of project landscapes and the integration process, with a view to improving the experience of both newcomers and the people responsible for orienting them, we performed a grounded theory study with 18 newcomers across 18 projects. We identified the main features that characterize a project landscape, together with key orientation aids and obstacles, and we theorize that there are three primary factors that impact the integration experience of newcomers: early experimentation, internalizing structures and cultures, and progress validation. Barthélémy Dagenais, Harold Ossher, Rachel K. E. Bellamy, Martin P. Robillard, Jacqueline de Vries |
ICSE (1) | 3 |
| 2010 | Flexible Modeling Tools (FlexiTools2010)abstractModeling tools are often not used for tasks during the software lifecycle for which they should be helpful; more free-from approaches, such as office tools and white boards, are frequently used instead. Why is this? What might be done to make modeling tools more suitable? What key research challenges must be overcome to achieve this? The goal of this workshop is to bring together people who understand the activities and needs of developers and other stakeholders throughout the software lifecycle, user interface design and tool infrastructure to explore these questions. Harold Ossher, André van der Hoek, Margaret-Anne D. Storey, John C. Grundy, Rachel K. E. Bellamy |
ICSE (2) | 5 |
| 2010 | Flexible modeling tools for pre-requirements analysis: conceptual architecture and research challengesabstractA serious tool gap exists at the start of the software lifecy-cle, before requirements formulation. Pre-requirements analysts gather information, organize it to gain insight, en-vision possible futures, and present insights and recom-mendations to stakeholders. They typically use office tools, which give great freedom, but no help with consistency management, change propagation, or information migration to downstream tools. Despite these downsides, office tools are still favored over modeling tools, which are constrain-ing and difficult to use. We introduce the notion of flexible modeling tools, which blend the advantages of office and modeling tools. We propose a conceptual architecture for such tools, and outline research challenges to be met in realizing them. We briefly describe the Business Insight Toolkit, a prototype tool embodying this architecture. Harold Ossher, Rachel K. E. Bellamy, Ian Simmonds, David Amid, Ateret Anaby-Tavor, Matthew Callery, Michael Desmond, Jacqueline de Vries, Amit Fisher, Sophia Krasikov |
OOPSLA | 2 |
| 2009 | An Empirical Study of Enterprise Conceptual Modeling
Ateret Anaby-Tavor, David Amid, Amit Fisher, Harold Ossher, Rachel K. E. Bellamy, Matthew Callery, Michael Desmond, Sophia Krasikov, Tova Roth, Ian Simmonds, Jacqueline de Vries |
ER | 5 |
| 2009 | An algorithm for identifying the abstract syntax of graph-based diagramsabstractDiagrams play a key role in the information systems domain. However to be meaningful, the diagrams are understood by interpreting visual cues in specific, conventionalized ways, termed conceptual models. One of the major pain points of conceptual models, specified as visual languages, is the inability to capture these visual languages effectively in conventional modeling tools. Instead, conceptual models are drawn using drawing tools and sometimes even by hand. We propose an automatic procedure to derive the syntactic building blocks of graph-based conceptual models. This high-level specification of the visual language can then serve as input for the automatic construction of syntax-aware diagram editors. Our aim is to achieve minimum effort on the part of the users when they eventually work with the graphical editor to produce a new diagram using the proposed syntax. Ateret Anaby-Tavor, David Amid, Amit Fisher, Harold Ossher, Rachel K. E. Bellamy, Matthew Callery, Michael Desmond, Sophia Krasikov, Tova Roth, Ian Simmonds, Jacqueline de Vries |
VL/HCC | 5 |
| 2008 | Using information scent to model the dynamic foraging behavior of programmers in maintenance tasksabstractIn recent years, the software engineering community has begun to study program navigation and tools to support it. Some of these navigation tools are very useful, but they lack a theoretical basis that could reduce the need for ad hoc tool building approaches by explaining what is fundamentally necessary in such tools. In this paper, we present PFIS (Programmer Flow by Information Scent), a model and algorithm of programmer navigation during software maintenance. We also describe an experimental study of expert programmers debugging real bugs described in real bug reports for a real Java application. We found that PFIS' performance was close to aggregated human decisions as to where to navigate, and was significantly better than individual programmers' decisions. Joseph Lawrance, Rachel K. E. Bellamy, Margaret M. Burnett, Kyle Rector |
CHI | 2 |
| 2008 | Can information foraging pick the fix? A field studyabstractPrevious findings have revealed the ability of information foraging to model or predict where developers will navigate within source code. However, the previous investigation did not consider whether the places developers went were the right places to go. In this paper, we present afield study in which we investigated over 200 open source bug reports and feature requests. We analyzed the textual similarity of these issues in relation to the source code, and determined what files developers had changed to fix these issues. Our results demonstrate that information scent can narrow down quite well where developers should make fixes, implying that future software navigation tools can predict the appropriate places to make fixes based solely on the contents of the issue and the source code. Joseph Lawrance, Rachel K. E. Bellamy, Margaret Bumett, Kyle Rector |
VL/HCC | 2 |
| 2007 | Evaluating an Automated Tool to Assist Evolutionary Document GenerationabstractWhile using how-to documents for guidance in performing computer-based tasks, users often run into problems due to inaccurate, out-of-date and incomplete documentation. These problems are often due to current documentation practices, which fail to keep how-to documents current, accurate, and complete. We believe that automated support for incremental update of how-to-documents, through the use of programming by demonstration and guided walkthrough techniques, is more effective than existing practice and produces documents that cause fewer problems for their users. In this paper, we present a study that evaluates this belief by comparing DocWizards, a tool utilizing these techniques, with a standard word processor. We show that more effective and efficient documentation can be generated by multiple authors using DocWizards in an incremental process, with effort comparable to that incurred using a traditional tool. Gahgene Gweon, Lawrence D. Bergman, Vittorio Castelli, Rachel K. E. Bellamy |
VL/HCC | 4 |
| 2007 | Scents in Programs: Does Information Foraging Theory Apply to Program Maintenance?abstractDuring maintenance, professional developers generate and test many hypotheses about program behavior, but they also spend much of their time navigating among classes and methods. Little is known, however, about how professional developers navigate source code and the extent to which their hypotheses relate to their navigation. A lack of understanding of these issues is a barrier to tools aiming to reduce the large fraction of time developers spend navigating source code. In this paper, we report on a study that makes use of information foraging theory to investigate how professional developers navigate source code during maintenance. Our results showed that information foraging theory was a significant predictor of the developers' maintenance behavior, and suggest how tools used during maintenance can build upon this result, simply by adding word analysis to their reasoning systems. Joseph Lawrance, Rachel K. E. Bellamy, Margaret M. Burnett |
VL/HCC | 2 |
| 1994 | What Does Pseudo-Code Do? A Psychological Analysis of the use of Pseudo-Code by Experienced ProgrammersabstractThe use of pseudo-code and pen and paper are prevalent within the task of programming. However, few studies examine the use of informal notations or the use of the paper medium. In this article, I offer a psychological analysis of the use of pseudo-code and pen and paper by experienced programmers. In particular, I investigate how such informal notations and the paper medium support the cognitively complex task of programming. The basis of the investigation is an analysis of the notes that programmers make during programming. These notes were collected from eight experienced programmers, who were all programming in different languages with different programming environments. Interviews and questionnaires were used as supplementary data. In the analysis based on these data, I describe the kinds of tasks done using pseudo-code and pen and paper, and I offer an account of why these tasks are done using these particular notations and this medium. This study suggests that programmers use pseudo-code and pen and paper to reduce the cognitive complexity of the programming task. Rachel K. E. Bellamy |
Hum. Comput. Interact. | 1 |
| 1992 | Re-Structuring the Programmer's Task
Rachel K. E. Bellamy, John M. Carroll 0001 |
Int. J. Man Mach. Stud. | 1 |
| 1990 | A view matcher for learning SmalltalkabstractThe View Matcher is a structured browser for Smalltalk/V. It presents a set of integrated and dynamic views of a running application, intended to coordinate and rationalize a programmer's early understanding of Smalltalk and its environment. We describe the system through two user scenarios involving exploration of the model-view-controller paradigm. John M. Carroll 0001, Janice A. Singer, Rachel K. E. Bellamy, Sherman R. Alpert |
CHI | 3 |
| 1990 | Smalltalk scaffolding: a case study of minimalist instructionabstractA curriculum was developed to introduce users to the Smalltalk object-oriented programming language. Applying the Minimalist model of instruction [3], we developed a set of example-based learning scenarios aimed at supporting real work, getting started fast, reasoning and improvising, coordinating system and text, supporting error recognition and recovery, and exploiting prior knowledge. We describe our initial curriculum design as well as the significant changes that have taken place as we have observed it in use. Mary Beth Rosson, John M. Carroll 0001, Rachel K. E. Bellamy |
CHI | 3 |
| 1990 | A psychology of programming for design
Rachel K. E. Bellamy |
INTERACT | 1 |
| 1990 | Redesign by design
Rachel K. E. Bellamy, John M. Carroll 0001 |
INTERACT | 1 |