Margaret M. Burnett

dblp:b/MMBurnett · also Margaret Burnett · DBLP profile ↗
← Back
139ranked-venue papers
17as first author
26since 2021 · last 2026
0000-0001-6536-7629ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 90 · 8 first-author · 15 since 2021Software engineering, systems software and programming languages · 34 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 "Fast, easy, simple"? SES-diverse transfer students' sociotechnical experiences registering for classes
abstract
Recruiting, retaining, and educating students in computing is a frequent research topic in CHI. However, students’ sociotechnical experiences of registering for classes are understudied—especially those of socioeconomic-diverse students. These experiences matter: research shows that registration problems bring long-term consequences to student successes. We investigate students’ socioeconomic status (SES) impact on registration experiences through three studies: a case study with education professionals using an emerging analytic method, SocioeconomicMag (SESMag); interviews with faculty/staff/students from 8 universities; and observations of 14 SES-diverse students registering for classes. Results showed: (1) 5 SES-inclusivity bugs which arose 30 times, 72% more often by lower-SES students than by higher-SES students. (2) 6/7 lower-SES students (but only 2/7 higher-SES students) expected downstream problems from the registration issues. (3) The risk-to-negative-outcomes rate was 3 times higher for lower-SES students. (4) The issues generalized across 8 universities and potentially to >700 other universities who use the same registration portal.
Alec Busteed, Jimena Noa Guevara, Lais Alexandra Castro, Dahana Moz-Ruiz, Iman Mokraoui, Prisha Velhal, Patricia Morreale, Anita Sarma, Margaret M. Burnett
CHI10
2026 "Over-the-Hood" AI Inclusivity Bugs and How 3 AI Product Teams Found and Fixed Them
abstract
While much research has shown the presence of AI’s “under-the-hood” biases (e.g., algorithmic, training data, etc.), what about “over-the-hood” inclusivity biases: barriers in user-facing AI products that disproportionately exclude users with certain problem-solving approaches? Recent research has begun to report the existence of such biases—but what do they look like, how prevalent are they, and how can developers find and fix them? To find out, we conducted a field study with 3 AI product teams, to investigate what kinds of AI inclusivity bugs exist uniquely in user-facing AI products, and whether/how AI product teams might harness an existing (non-AI-oriented) inclusive design method to find and fix them. The teams’ work revealed 83 instances of 6 AI inclusivity bug types unique to user-facing AI products, their fixes covering 47 bug instances, and a new GenderMag inclusive design method variant, GenderMag-for-AI, that is especially effective at detecting AI inclusivity bugs when the AI’s output is not necessarily believed.
Andrew Anderson 0002, Fatima A. Moussaoui, Jimena Noa Guevara, Md Montaser Hamid, Margaret M. Burnett
IUI5
2026 Inclusive Design of AI's Explanations: Just for Those Previously Left Out?
abstract
Abstract Motivations . Explainable AI (XAI) systems aim to improve users’ understanding of AI, but XAI research has shown that many XAI explanations serve some users well while failing others. In non-AI systems, software practitioners have used inclusive design approaches to address similar problems, sometimes creating “curb-cut” improvements that benefit both underserved users and everyone else. This raises the possibility that inclusive design approaches can bring similar curb-cut improvements to AI explanations. Objectives . Our objective was to investigate possible curb-cut effects of inclusivity-driven fixes an AI product team made using an inclusive design approach (GenderMag) to improve their XAI prototype. Methods . We ran a between-subject study with 69 participants who had no formal AI background. 34 participants used the original version of the XAI prototype and the rest used the version with the AI team’s inclusivity fixes. We then compared the two groups’ mental model concepts scores and prediction accuracy, and the two prototypes’ inclusivity. Results . Our investigation produced four main results. First, the AI team’s inclusivity fixes were overall effective, resulting in overall better conceptual mental models with the new prototype. Further (second), the AI team’s inclusivity fixes were particularly beneficial to the underserved population’s conceptual mental models—which, together with the first result, constitutes a curb-cut effect. However (third), the inclusivity fixes did not improve participants’ prediction accuracy scores. Instead, it appears to have harmed them overall—a “curb-fence” effect (opposite of a curb-cut effect). Finally (fourth), the AI team’s fixes improved equity, reducing the gender gap by 45%.
Md Montaser Hamid, Fatima A. Moussaoui, Jimena Noa Guevara, Andrew Anderson 0002, Puja Agarwal, Jonathan Dodge, Margaret M. Burnett
ACM Trans. Interact. Intell. Syst.7
2025 How We Did It: Integrating Inclusive Design across the Undergraduate Computer Science Curriculum
abstract
Inclusive design appears rarely, if at all, in most undergraduate computer science (CS) curricula. As a result, many CS students graduate without knowing how to apply inclusive design to the software they build, and go on to careers that perpetuate the prolif- eration of software that excludes communities of users. Our panel of CS faculty will explain how we have been working to address this problem. For the past several years, we have been integrating bits of inclusive design in multiple courses in CS undergraduate programs, which has had very positive impacts on students' ratings of their instructors, students' ratings of the education climate, and students' retention. The panel's content will be mostly concrete examples of how we are doing this, so that attendees can leave with an in-the-trenches understanding of what this looks like for CS faculty across specialization areas and classes. We also show how it can be used in a department's BPC Plan and point to resources on the CRA's BPCnet Activity Library and on OERcommons, to enable interested faculty to go forward with this approach in their own classes and departments.
Patricia Morreale, Margaret M. Burnett, Kyle J. Harms, Daehan Kwak
SIGCSE (2)2
2025 Measuring SES-related traits relating to technology usage: Two validated surveys
abstract
Abstract Software producers are now recognizing the importance of improving their products’ suitability for diverse populations, but little attention has been given to measurements to shed light on products’ suitability to individuals below the median s ocio e conomic s tatus (SES)—who, by definition, make up half the population. To enable software practitioners to attend to both lower- and higher-SES individuals, this paper provides two new surveys that together can facilitate measuring how well a software product serves socioeconomically diverse populations. The first survey (SES-Subjective) is who-oriented: it measures who their potential or current users are in terms of their subjective SES (perceptions of their SES). The second survey (SES-Facets) is why-oriented: it collects individuals’ values for an evidence-based set of facet values (individual traits) that (1) statistically differ by SES and (2) affect how an individual works and problem-solves with software products. The surveys’ design goal is worldwide applicability, but as a first step, here we empirically validated both these surveys with deployments at University A and University B (464 and 522 responses, respectively), which showed reliability of both the surveys in a US context. Our results also statistically agree with both ground truth data on respondents’ socioeconomic statuses and with predictions from foundational literature. Finally, we explain how the pair of surveys can be uniquely actionable by software practitioners, such as in requirements gathering, debugging, quality assurance activities, maintenance activities, and fulfilling legal reporting requirements such as those being drafted by various governments for AI-powered software.
Chimdi Chikezie, Pannapat Chanpaisaeng, Puja Agarwal, Bhavika Madhwani, Rudrajit Choudhuri, Andrew Anderson 0002, Prisha Velhal, Patricia Morreale, Christopher Bogart, Anita Sarma, Margaret M. Burnett
Empir. Softw. Eng.12
2025 Intersectional HCI on a Budget: An Analytical Approach Powered by Types
abstract
Intersectional HCI recognizes that humans' interconnected social identities shape their experiences with technology. However, intersectional HCI requires extensive resources, such as access to intersectional populations, which many HCI practitioners may lack. For these practitioners, we present an analytical approach to bring intersectional lenses to HCI practices. The approach uses types—not at the level of identities, but at the level of personal traits drawn from foundational research. We first formally prove that certain analytical methods for detecting inclusivity issues can be meaningfully composed to provide equitable consideration of typically overlooked populations; then present four design use-cases to illustrate what the approach brings to HCI practices; and then empirically investigated one of the four use-cases with 24 HCI participants. Results show that practitioners using the compositional approach detected even more intersectional inclusivity problems than those using a complementary intersectional approach.
Abrar Fallatah, Md Montaser Hamid, Fatima A. Moussaoui, Chimdi Chikezie, Martin Erwig, Christopher Bogart, Anita Sarma, Margaret M. Burnett
Int. J. Hum. Comput. Interact.8
2024 The Matchmaker Inclusive Design Curriculum: A Faculty-Enabling Curriculum to Teach Inclusive Design Throughout Undergraduate CS
abstract
Despite efforts to raise awareness of societal and ethical issues in CS education, research shows students often do not act upon their new awareness (Problem 1). One such issue, well-established by HCI research, is that much of technology contains barriers impacting numerous populations—such as minoritized genders, races, ethnicities, and more. HCI has inclusive design methods that help—but these skills are rarely taught, even in HCI classes (Problem 2). To address Problems 1 and 2, we created the Matchmaker Curriculum to pair CS faculty—including non-HCI faculty—with inclusive design elements to allow for inclusive design skill-building throughout their CS program. We present the curriculum and a field study, in which we followed 18 faculty along their journey. The results show how the Matchmaker Curriculum equipped 88% of these faculty with enough inclusive design teaching knowledge to successfully embed actionable inclusive design skill-building into 13 CS courses.
Rosalinda Garcia, Patricia Morreale, Gail Verdi, Heather Garcia, Geraldine Jimena Noa, Spencer P. Madsen, Maria Jesus Alzugaray-Orellana, Elizabeth Li, Margaret M. Burnett
CHI9
2024 Debugging for Inclusivity in Online CS Courseware: Does it Work?
abstract
Online computer science (CS) courses have broadened access to CS education, yet inclusivity barriers persist for minoritized groups in these courses. One problem that recent research has shown is that often inclusivity biases (“inclusivity bugs”) lurk within the course materials themselves, disproportionately disadvantaging minoritized students. To address this issue, we investigated how a faculty member can use AID—an Automated Inclusivity Detector tool—to remove such inclusivity bugs from a large online CS1 (Intro CS) course and what is the impact of the resulting inclusivity fixes on the students’ experiences. To enable this evaluation, we first needed to (Bugs): investigate inclusivity challenges students face in 5 online CS courses; (Build): build decision rules to capture these challenges in courseware (“inclusivity bugs”) and implement them in the AID tool; (Faculty): investigate how the faculty member followed up on the inclusivity bugs that AID reported; and (Students): investigate how the faculty member’s changes impacted students’ experiences via a before-vs-after qualitative study with CS students. Our results from (Bugs) revealed 39 inclusivity challenges spanning courseware components from the syllabus to assignments. After implementing the rules in the tool (Build), our results from (Faculty) revealed how the faculty member treated AID more as a “peer” than an authority in deciding whether and how to fix the bugs. Finally, the study results with (Students) revealed that students found the after-fix courseware more approachable - feeling less overwhelmed and more in control in contrast to the before-fix version where they constantly felt overwhelmed, often seeking external assistance to understand course content.
Amreeta Chatterjee, Rudrajit Choudhuri, Mrinmoy Sarkar, Soumiki Chattopadhyay, Dylan Liu, Samarendra Hedaoo, Margaret M. Burnett, Anita Sarma
ICER (1)7
2024 Beyond "Awareness": If We Teach Inclusive Design, Will Students Act On It?
abstract
Motivation: Many university CS programs have begun teaching various types of CS-related societal issues using approaches such as ethics, Responsible CS, inclusive design, and more. However, some recent research suggests that, although these programs have been able to teach awareness, students often fail to act upon this awareness. To address this problem, University X's CS program tried an unusual approach—integrating hands-on inclusive design skills in small ways across all four years of the CS major. But did it work? That is, did the students who experienced this change across the major actually build more inclusive technology than the students who did not experience it?
Rosalinda Garcia, Patricia Morreale, Pankati Patel, Jimena Noa Guevara, Dahana Moz-Ruiz, Sabyatha Sathish Kumar, Prisha Velhal, Alec Busteed, Margaret M. Burnett
ICER (1)9
2024 From Workshops to Classrooms: Faculty Experiences with Implementing Inclusive Design Principles
abstract
Computer science (CS) and information technology (IT) curricula are grounded in theoretical and technical skills. Topics like equity and inclusive design are rarely found in mainstream student studies. This results in graduates with outdated practices and limitations in software development. A research project was conducted to educate the faculty to integrate inclusive software design into the CS undergraduate curriculum. The objective is to produce graduates with the ability to develop inclusive software.
Pankati Patel, Dahana Moz-Ruiz, Rosalinda Garcia, Amreeta Chatterjee, Patricia Morreale, Margaret M. Burnett
SIGCSE (1)6
2024 Measuring User Experience Inclusivity in Human-AI Interaction via Five User Problem-Solving Styles
abstract
Motivations : Recent research has emerged on generally how to improve AI products’ human-AI interaction (HAI) user experience (UX), but relatively little is known about HAI-UX inclusivity. For example, what kinds of users are supported, and who are left out? What product changes would make it more inclusive? Objectives : To help fill this gap, we present an approach to measuring what kinds of diverse users an AI product leaves out and how to act upon that knowledge. To bring actionability to the results, the approach focuses on users’ problem-solving diversity. Thus, our specific objectives were (1) to show how the measure can reveal which participants with diverse problem-solving styles were left behind in a set of AI products and (2) to relate participants’ problem-solving diversity to their demographic diversity, specifically gender and age. Methods : We performed 18 experiments, discarding two that failed manipulation checks. Each experiment was a 2 \(\times\) 2 factorial experiment with online participants, comparing two AI products: one deliberately violating 1 of 18 HAI guidelines and the other applying the same guideline. For our first objective, we used our measure to analyze how much each AI product gained/lost HAI-UX inclusivity compared to its counterpart, where inclusivity meant supportiveness to participants with particular problem-solving styles. For our second objective, we analyzed how participants’ problem-solving styles aligned with their gender identities and ages. Results and Implications : Participants’ diverse problem-solving styles revealed six types of inclusivity results: (1) the AI products that followed an HAI guideline were almost always more inclusive across diversity of problem-solving styles than the products that did not follow that guideline—but “who” got most of the inclusivity varied widely by guideline and by problem-solving style; (2) when an AI product had risk implications, four variables’ values varied in tandem: participants’ feelings of control, their (lack of) suspicion, their trust in the product, and their certainty while using the product; (3) the more control an AI product offered users, the more inclusive it was; (4) whether an AI product was learning from “my” data or other people’s affected how inclusive that product was; (5) participants’ problem-solving styles skewed differently by gender and age group; and (6) almost all of the results suggested actions that HAI practitioners could take to improve their products’ inclusivity further. Together, these results suggest that a key to improving the demographic inclusivity of an AI product (e.g., across a wide range of genders, ages) can often be obtained by improving the product’s support of diverse problem-solving styles.
Andrew Anderson 0002, Jimena Noa Guevara, Fatima A. Moussaoui, Tianyi Li 0008, Mihaela Vorvoreanu, Margaret M. Burnett
ACM Trans. Interact. Intell. Syst.6
2023 Implementing Inclusive Software Design in the CS Curriculum
Pankati Patel, Jean Chu, Yulia Kumar, Daehan Kwak, Patricia Morreale, Rosalinda Garcia, Margaret M. Burnett
SIGCSE (2)7
2023 Embedding Equitable Design in the CS Computing Curricula
abstract
Computer science (CS) students' curricula is heavily focused on technical skills, and CS ethics, usability, equity, and people/society considerations are not well-integrated into the CS curriculum. If these topics are introduced, they are disconnected from the core courses. As a result, students do not learn to incorporate inclusive practices into their software designs. Thus, students create software through the perspective of a computer scientist - when the important perspective is that of the intended users. As a result, students entering the workforce are inclined to design software that is non-inclusive. We propose the integration of inclusive design in the undergraduate curriculum will result in students creating inclusive software. This research, based on the foundations of inclusive design methods, investigates a new approach to teaching CS. Inclusive software design is embedded into computing courses for all four years of the undergraduate CS curriculum. This new approach is "minimally invasive", occupying very little classroom time, instead it is integrated into the course work that is already assigned. With this work, we hope to answer the following questions: (1) Will this new approach improve students' ability to design inclusive software? (2) Will this approach create an inclusive climate among peers? (3) Will it affect student's success or lack thereof? (4) How and to what extent is this embedded inclusive design curriculum feasible to use?
Pankati Patel, Patricia Morreale, Yulia Kumar, Daehan Kwak, Jean Chu, Rose Garcia, Margaret M. Burnett
SIGCSE (2)7
2023 Getting Outside the Bug Boxes (Keynote)
abstract
Sometimes, we humans find ourselves a bit slow to abandon the comfort of sitting “inside the box”, and this can detract from our ability to innovate. In this talk, I’ll share some outside-the-box perspectives, gleaned from decades of software engineering work, on boxes I’ve seen when thinking about bugs — from failures to faults, from finding to fixing, and from traditional to very non-traditional notions of “what counts” as a bug. I’ll consider the intellectually freeing perspectives that can come from moving outside the “mechanisms” box to policies; the enhancement to applicability from moving outside sub-sub-area boxes to the whole software lifecycle; the differences revealed when moving outside the “typical developer” box to diverse humans; and the plethora of possibilities arising from moving outside the “buggy code” box to a wide range of bug types.
Margaret M. Burnett
ESEC/SIGSOFT FSE1
2023 "Regular" CS × Inclusive Design = Smarter Students and Greater Diversity
abstract
What if “regular” Computer Science (CS) faculty each taught elements of inclusive design in “regular” CS courses across an undergraduate curriculum? Would it affect the CS program's climate and inclusiveness to diverse students? Would it improve retention? Would students learn less CS? Would they actually learn any inclusive design? To answer these questions, we conducted a year-long Action Research investigation, in which 13 CS faculty integrated elements of inclusive design into 44 CS/IT offerings across a 4-year curriculum. The 613 affected students’ educational work products, grades, and/or climate questionnaire responses revealed significant improvements in students’ course outcomes (higher course grades and fewer course fails/incompletes/withdrawals), especially for marginalized groups; revealed that most students did learn and apply inclusive design concepts to their CS activities; and revealed that inclusion and teamwork in the courses significantly improved. These results suggest a new pathway for significantly improving students’ retention, their knowledge and usage of inclusive design, and their experiences across CS education—for marginalized groups and for all students.
Rosalinda Garcia, Patricia Morreale, Lara Letaw, Amreeta Chatterjee, Pankati Patel, Sarah Yang, Isaac Tijerina Escobar, Geraldine Jimena Noa, Margaret M. Burnett
ACM Trans. Comput. Educ.9
2022 Inclusivity Bugs in Online Courseware: A Field Study
abstract
Motivation: Although asynchronous online CS courses have enabled more diverse populations to access CS higher education, research shows that online CS-ed is far from inclusive, with women and other underrepresented groups continuing to face inclusion gaps. Worse, diversity/inclusion research in CS-ed has largely overlooked the online courseware—the web pages and course materials that populate the online learning platforms—that constitute asynchronous online CS-ed’s only mechanism of course delivery.
Amreeta Chatterjee, Lara Letaw, Rosalinda Garcia, Doshna Umma Reddy, Rudrajit Choudhuri, Sabyatha Sathish Kumar, Patricia Morreale, Anita Sarma, Margaret M. Burnett
ICER (1)9
2022 How Do People Rank Multiple Mutant Agents?
abstract
Faced with several AI-powered sequential decision-making systems, how might someone choose on which to rely? For example, imagine car buyer Blair shopping for a self-driving car, or developer Dillon trying to choose an appropriate ML model to use in their application. Their first choice might be infeasible (i.e., too expensive in money or execution time), so they may need to select their second or third choice. To address this question, this paper presents: 1) Explanation Resolution, a quantifiable direct measurement concept; 2) a new XAI empirical task to measure explanations: “the Ranking Task”; and 3) a new strategy for inducing controllable agent variations—Mutant Agent Generation. In support of those main contributions, it also presents 4) novel explanations for sequential decision-making agents; 5) an adaptation to the AAR/AI assessment process; and 6) a qualitative study around these devices with 10 participants to investigate how they performed the Ranking Task on our mutant agents, using our explanations, and structured by AAR/AI. From an XAI researcher perspective, just as mutation testing can be applied to any code, mutant agent generation can be applied to essentially any neural network for which one wants to evaluate an assessment process or explanation type. As to an XAI user’s perspective, the participants ranked the agents well overall, but showed the importance of high explanation resolution for close differences between agents. The participants also revealed the importance of supporting a wide diversity of explanation diets and agent “test selection” strategies.
Jonathan Dodge, Andrew Anderson 0002, Matthew L. Olson, Rupika Dikkala, Margaret M. Burnett
IUI5
2022 Finding AI's Faults with AAR/AI: An Empirical Study
abstract
Would you allow an AI agent to make decisions on your behalf? If the answer is “not always,” the next question becomes “in what circumstances”? Answering this question requires human users to be able to assess an AI agent—and not just with overall pass/fail assessments or statistics. Here users need to be able to localize an agent’s bugs so that they can determine when they are willing to rely on the agent and when they are not. After-Action Review for AI (AAR/AI), a new AI assessment process for integration with Explainable AI systems, aims to support human users in this endeavor, and in this article we empirically investigate AAR/AI’s effectiveness with domain-knowledgeable users. Our results show that AAR/AI participants not only located significantly more bugs than non-AAR/AI participants did (i.e., showed greater recall) but also located them more precisely (i.e., with greater precision). In fact, AAR/AI participants outperformed non-AAR/AI participants on every bug and were, on average, almost six times as likely as non-AAR/AI participants to find any particular bug. Finally, evidence suggests that incorporating labeling into the AAR/AI process may encourage domain-knowledgeable users to abstract above individual instances of bugs; we hypothesize that doing so may have contributed further to AAR/AI participants’ effectiveness.
Roli Khanna, Jonathan Dodge, Andrew Anderson 0002, Rupika Dikkala, Jed Irvine, Zeyad Shureih, Kin-Ho Lam, Caleb R. Matthews, Zhengxian Lin, Minsuk Kahng, Alan Fern, Margaret M. Burnett
ACM Trans. Interact. Intell. Syst.12
2022 How Gender-Biased Tools Shape Newcomer Experiences in OSS Projects
abstract
Previous research has revealed that newcomer women are disproportionately affected by gender-biased barriers in open source software (OSS) projects. However, this research has focused mainly on social/cultural factors, neglecting the software tools and infrastructure. To shed light on how OSS tools and infrastructure might factor into OSS barriers to entry, we conducted two studies: (1) a field study with five teams of software professionals, who worked through five use cases to analyze the tools and infrastructure used in their OSS projects; and (2) a diary study with 22 newcomers (9 women and 13 men) to investigate whether the barriers matched the ones identified by the software professionals. The field study produced a bleak result: software professionals found gender biases in 73 percent of all the newcomer barriers they identified. Further, the diary study confirmed these results: Women newcomers encountered gender biases in 63 percent of barriers they faced. Fortunately, many kinds of barriers and biases revealed in these studies could potentially be ameliorated through changes to the OSS software environments and tools.
Hema Susmita Padala, Christopher J. Mendez, Felipe Fronchetti, Igor Steinmacher, Zoe Steine-Hanson, Claudia Hilderbrand, Amber Horvath, Charles Hill 0001, Logan Simpson, Margaret M. Burnett, Marco Aurélio Gerosa, Anita Sarma
IEEE Trans. Software Eng.10
2021 Changing the Online Climate via the Online Students: Effects of Three Curricular Interventions on Online CS Students' Inclusivity
abstract
Motivation: Although CS Education researchers and practitioners have found ways to improve CS classroom inclusivity, few researchers have considered inclusivity of online CS education. We are interested in two such improvements in online CS education—besides being inclusive to each other, online CS students also need to be able to create inclusive technology. Objectives: We have begun developing a new approach that we term “embedded inclusive design” to address both of these goals. The essence of the approach is to integrate elements of inclusive design education into mainstream CS coursework. This paper presents three curricular interventions we have developed in this approach and empirically investigates their efficacy in online CS post-baccalaureate education. Our research questions were: How do these three curricular interventions affect (RQ1) the climate among online CS students and (RQ2) how online CS students honor the diversity of their users in the tech they create? Method: To answer these research questions, we implemented the curricular interventions in four asynchronous online CS classes across two CS courses within Oregon State University’s Ecampus and conducted an action research study to investigate the impacts. Results: Online CS students who experienced these interventions reported feeling more included in the major than they had before, reported positive impacts on their team dynamics, increased their interest in accommodating diverse users, and created more inclusive technology designs than they had before. Discussion: These results provide encouraging evidence that embedding elements of inclusive design into mainstream CS coursework, via the interventions presented here, can increase both online CS students’ inclusivity toward one another and the inclusivity of the technology these future CS practitioners create.
Lara Letaw, Rosalinda Garcia, Heather Garcia, Christopher Perdriau, Margaret M. Burnett
ICER5
2021 AID: An automated detector for gender-inclusivity bugs in OSS project pages
abstract
The tools and infrastructure used in tech, including Open Source Software (OSS), can embed "inclusivity bugs"- features that disproportionately disadvantage particular groups of contributors. To see whether OSS developers have existing practices to ward off such bugs, we surveyed 266 OSS developers. Our results show that a majority (77%) of developers do not use any inclusivity practices, and 92% of respondents cited a lack of concrete resources to enable them to do so. To help fill this gap, this paper introduces AID, a tool that automates the GenderMag method to systematically find gender-inclusivity bugs in software. We then present the results of the tool's evaluation on 20 GitHub projects. The tool achieved precision of 0.69, recall of 0.92, an F-measure of 0.79 and even captured some inclusivity bugs that human GenderMag teams missed.
Amreeta Chatterjee, Mariam Guizani, Catherine Stevens, Jillian Emard, Mary Evelyn May, Margaret M. Burnett, Iftekhar Ahmed 0001, Anita Sarma
ICSE6
2021 Artificial Intelligence versus End-User Development: A Panel on What Are the Tradeoffs in Daily Automations?
Fabio Paternò, Margaret M. Burnett, Gerhard Fischer, Maristella Matera, Brad A. Myers, Albrecht Schmidt 0001
INTERACT (5)2
2021 Towards User-Centric Robot Furniture Arrangement
abstract
Imbuing furniture with robot properties reduces the physical labor and time needed for arranging spaces, such as homes, classrooms, and offices. Outsourcing labor tasks to robot furniture requires users’ involvement with functional user interfaces. We performed a user study on multi-robot furniture and added additional features based on the study results. The study involved 12 participants rearranging multiple non-robotic and robotic chairs (ChairBots). Results from the video and interview analysis revealed five high-level features missing in the original ChairBot: dual screen-based user interface, the ability to save and to set arrangements, the ability to move in multi-robot formations, the ability to snap to angles/gridlines, and higher movement precision. The improved system allows users to control multiple furniture robots, both locally and remotely. Such improvement sets the baseline functionalities of robot furniture arrangement systems while extending the potential utilization of established robotic chairs.
Abrar Fallatah, Brett Stoddard, Margaret M. Burnett, Heather Knight
RO-MAN3
2021 After-Action Review for AI (AAR/AI)
abstract
Explainable AI is growing in importance as AI pervades modern society, but few have studied how explainable AI can directly support people trying to assess an AI agent. Without a rigorous process, people may approach assessment in ad hoc ways—leading to the possibility of wide variations in assessment of the same agent due only to variations in their processes. AAR, or After-Action Review, is a method some military organizations use to assess human agents, and it has been validated in many domains. Drawing upon this strategy, we derived an After-Action Review for AI (AAR/AI), to organize ways people assess reinforcement learning agents in a sequential decision-making environment. We then investigated what AAR/AI brought to human assessors in two qualitative studies. The first investigated AAR/AI to gather formative information, and the second built upon the results, and also varied the type of explanation (model-free vs. model-based) used in the AAR/AI process. Among the results were the following: (1) participants reporting that AAR/AI helped to organize their thoughts and think logically about the agent, (2) AAR/AI encouraged participants to reason about the agent from a wide range of perspectives , and (3) participants were able to leverage AAR/AI with the model-based explanations to falsify the agent’s predictions.
Jonathan Dodge, Roli Khanna, Jed Irvine, Kin-Ho Lam, Theresa Mai, Zhengxian Lin, Nicholas Kiddle, Evan Newman, Andrew Anderson 0002, Sai Raja, Caleb R. Matthews, Christopher Perdriau, Margaret M. Burnett, Alan Fern
ACM Trans. Interact. Intell. Syst.13
2021 The Shoutcasters, the Game Enthusiasts, and the AI: Foraging for Explanations of Real-time Strategy Players
abstract
Assessing and understanding intelligent agents is a difficult task for users who lack an AI background. “Explainable AI” (XAI) aims to address this problem, but what should be in an explanation? One route toward answering this question is to turn to theories of how humans try to obtain information they seek. Information Foraging Theory (IFT) is one such theory. In this article, we present a series of studies 1 using IFT: the first investigates how expert explainers supply explanations in the RTS domain, the second investigates what explanations domain experts demand from agents in the RTS domain, and the last focuses on how both populations try to explain a state-of-the-art AI. Our results show that RTS environments like StarCraft offer so many options that change so rapidly, foraging tends to be very costly. Ways foragers attempted to manage such costs included “satisficing” approaches to reduce their cognitive load, such as focusing more on What information than on Why information, strategic use of language to communicate a lot of nuanced information in a few words, and optimizing their environment when possible to make their most valuable information patches readily available. Further, when a real AI entered the picture, even very experienced domain experts had difficulty understanding and judging some of the AI’s unconventional behaviors. Finally, our results reveal ways Information Foraging Theory can inform future XAI interactive explanation environments, and also how XAI can inform IFT.
Sean Penney, Jonathan Dodge, Andrew Anderson 0002, Claudia Hilderbrand, Logan Simpson, Margaret M. Burnett
ACM Trans. Interact. Intell. Syst.6
2021 Version Control Systems: An Information Foraging Perspective
abstract
Version Control Systems (VCS) are an important source of information for developers. This calls for a principled understanding of developers' information seeking in VCS-both for improving existing tools and for understanding requirements for new tools. Our prior work investigated empirically how and why developers seek information in VCS: in this paper, we complement and enrich our prior findings by reanalyzing the data via a theory's lens. Using the lens of Information Foraging Theory (IFT), we present new insights not revealed by the prior empirical work. First, while looking for specific information, participants' foraging behaviors were consistent with other foraging situations in SE; therefore, prior research on IFT-based SE tool design can be leveraged for VCS. Second, in change awareness foraging, participants consumed similar diets, but in subtly different ways than in other situations; this calls for further investigations into change awareness foraging. Third, while committing changes, participants attempted to enable future foragers, but the competing needs of different foraging situations led to tensions that participants failed to balance: this opens up a new avenue for research at the intersection of IFT and SE, namely, creating forageable information. Finally, the results of using an IFT lens on these data provides some evidence as to IFT's scoping and utility for the version control domain.
Sruti Srinivasa Ragavan, Mihai Codoban, David Piorkowski, Danny Dig, Margaret M. Burnett
IEEE Trans. Software Eng.5
2020 Doing Inclusive Design: From GenderMag in the Trenches to Inclusive Mag in the Research Lab
abstract
How can user interface and user experience (UI/UX) professionals assess whether their software supports diverse users? And if they find problems, how can they fix them? We begin this keynote address with a summary of GenderMag, a systematic inspection method for finding and fixing "gender inclusivity bugs---biases against different genders in software interfaces and workflows. We then show what UI/UX professionals are doing with it in the real world, from their bias finds & fixes to their practices & pitfalls in using it. Finally, we present InclusiveMag, a meta-method that can be used by HCI researchers to generate systematic inclusiveness methods for other dimensions of diversity.
Margaret M. Burnett
AVI1
2020 Engineering gender-inclusivity into software: ten teams' tales from the trenches
abstract
Although the need for gender-inclusivity in software is gaining attention among SE researchers and SE practitioners, and at least one method (GenderMag) has been published to help, little has been reported on how to make such methods work in real-world settings. Real-world teams are ever-mindful of the practicalities of adding new methods on top of their existing processes. For example, how can they keep the time costs viable? How can they maximize impacts of using it? What about controversies that can arise in talking about gender? To find out how software teams "in the trenches" handle these and similar questions, we collected the GenderMag-based processes of 10 real-world software teams---more than 50 people---for periods ranging from 5 months to 3.5 years. We present these teams' insights and experiences in the form of 9 practices, 2 potential pitfalls, and 2 open issues, so as to provide their insights to other real-world software teams trying to engineer gender-inclusivity into their software products.
Claudia Hilderbrand, Christopher Perdriau, Lara Letaw, Jillian Emard, Zoe Steine-Hanson, Margaret M. Burnett, Anita Sarma
ICSE6
2020 Keeping it "organized and logical": after-action review for AI (AAR/AI)
abstract
Explainable AI (XAI) is growing in importance as AI pervades modern society, but few have studied how XAI can directly support people trying to assess an AI agent. Without a rigorous process, people may approach assessment in ad hoc ways---leading to the possibility of wide variations in assessment of the same agent due only to variations in their processes. AAR, or After-Action Review, is a method some military organizations use to assess human agents, and it has been validated in many domains. Drawing upon this strategy, we derived an AAR for AI, to organize ways people assess reinforcement learning (RL) agents in a sequential decision-making environment. The results of our qualitative study revealed several strengths and weaknesses of the AAR/AI process and the explanations embedded within it.
Theresa Mai, Roli Khanna, Jonathan Dodge, Jed Irvine, Kin-Ho Lam, Zhengxian Lin, Nicholas Kiddle, Evan Newman, Sai Raja, Caleb R. Matthews, Christopher Perdriau, Margaret M. Burnett, Alan Fern
IUI12
2020 Mental Models of Mere Mortals with Explanations of Reinforcement Learning
abstract
How should reinforcement learning (RL) agents explain themselves to humans not trained in AI? To gain insights into this question, we conducted a 124-participant, four-treatment experiment to compare participants’ mental models of an RL agent in the context of a simple Real-Time Strategy (RTS) game. The four treatments isolated two types of explanations vs. neither vs. both together. The two types of explanations were as follows: (1) saliency maps (an “Input Intelligibility Type” that explains the AI’s focus of attention) and (2) reward-decomposition bars (an “Output Intelligibility Type” that explains the AI’s predictions of future types of rewards). Our results show that a combined explanation that included saliency and reward bars was needed to achieve a statistically significant difference in participants’ mental model scores over the no-explanation treatment. However, this combined explanation was far from a panacea: It exacted disproportionately high cognitive loads from the participants who received the combined explanation. Further, in some situations, participants who saw both explanations predicted the agent’s next action worse than all other treatments’ participants.
Andrew Anderson 0002, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Matthew L. Olson, Alan Fern, Margaret M. Burnett
ACM Trans. Interact. Intell. Syst.10
2020 Special Issue on Highlights of ACM Intelligent User Interface (IUI) 2018
abstract
research-article Share on Special Issue on Highlights of ACM Intelligent User Interface (IUI) 2018 Authors: Mark Billinghurst School of ITMS, University of South Australia, Adelaide, South Australia, Australia School of ITMS, University of South Australia, Adelaide, South Australia, AustraliaView Profile , Margaret Burnett School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, Oregon, USA School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, Oregon, USAView Profile , Aaron Quigley School of Computer Science, University of St. Andrews, St. Andrews, Scotland, United Kingdom School of Computer Science, University of St. Andrews, St. Andrews, Scotland, United KingdomView Profile Authors Info & Claims ACM Transactions on Interactive Intelligent SystemsVolume 10Issue 1March 2020 Article No.: 1pp 1–3https://doi.org/10.1145/3357206Published:12 October 2019Publication History 0citation192DownloadsMetricsTotal Citations0Total Downloads192Last 12 Months27Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Mark Billinghurst, Margaret M. Burnett, Aaron J. Quigley
ACM Trans. Interact. Intell. Syst.2
2019 From Gender Biases to Gender-Inclusive Design: An Empirical Investigation
abstract
In recent years, research has revealed gender biases in numerous software products. But although some researchers have found ways to improve gender participation in specific software projects, general methods focus mainly on detecting gender biases -- not fixing them. To help fill this gap, we investigated whether the GenderMag bias detection method can lead directly to designs with fewer gender biases. In our 3-step investigation, two HCI researchers analyzed an industrial software product using GenderMag; we derived design changes to the product using the biases they found; and ran an empirical study of participants using the original product versus the new version. The results showed that using the method in this way did improve the software's inclusiveness: women succeeded more often in the new version than in the original; men's success rates improved too; and the gender gap entirely disappeared.
Mihaela Vorvoreanu, Lingyi Zhang, Yun-Han Huang, Claudia Hilderbrand, Zoe Steine-Hanson, Margaret M. Burnett
CHI6
2019 Explaining Reinforcement Learning to Mere Mortals: An Empirical Study
abstract
We present a user study to investigate the impact of explanations on non-experts? understanding of reinforcement learning (RL) agents. We investigate both a common RL visualization, saliency maps (the focus of attention), and a more recent explanation type, reward-decomposition bars (predictions of future types of rewards). We designed a 124 participant, four-treatment experiment to compare participants? mental models of an RL agent in a simple Real-Time Strategy (RTS) game. Our results show that the combination of both saliency and reward bars were needed to achieve a statistically significant improvement in mental model score over the control. In addition, our qualitative analysis of the data reveals a number of effects for further study.
Andrew Anderson 0002, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Alan Fern, Margaret M. Burnett
IJCAI9
2019 From GenderMag to InclusiveMag: An Inclusive Design Meta-Method
abstract
How can software practitioners assess whether their software supports diverse users? Although there are empirical processes that can be used to find “inclusivity bugs” piecemeal, what is often needed is a systematic inspection method to assess software's support for diverse populations. To help fill this gap, this paper introduces InclusiveMag, a generalization of GenderMag that can be used to generate systematic inclusiveness methods for a particular dimension of diversity. We then present a multicase study covering eight diversity dimensions, of eight teams' experiences applying InclusiveMag to eight under-served populations and their “mainstream” counterparts.
Christopher J. Mendez, Lara Letaw, Margaret M. Burnett, Simone Stumpf, Anita Sarma, Claudia Hilderbrand
VL/HCC3
2018 How the Experts Do It: Assessing and Explaining Agent Behaviors in Real-Time Strategy Games
abstract
How should an AI-based explanation system explain an agent's complex behavior to ordinary end users who have no background in AI? Answering this question is an active research area, for if an AI-based explanation system could effectively explain intelligent agents' behavior, it could enable the end users to understand, assess, and appropriately trust (or distrust) the agents attempting to help them. To provide insights into this question, we turned to human expert explainers in the real-time strategy domain --"shoutcasters"-- to understand (1) how they foraged in an evolving strategy game in real time, (2) how they assessed the players' behaviors, and (3) how they constructed pertinent and timely explanations out of their insights and delivered them to their audience. The results provided insights into shoutcasters' foraging strategies for gleaning information necessary to assess and explain the players; a characterization of the types of implicit questions shoutcasters answered; and implications for creating explanations by using the patterns and abstraction levels these human experts revealed.
Jonathan Dodge, Sean Penney, Claudia Hilderbrand, Andrew Anderson 0002, Margaret M. Burnett
CHI5
2018 Pedagogical Content Knowledge for Teaching Inclusive Design
abstract
Inclusive design is important in today's software industry, but there is little research about how to teach it. In collaboration with 9 teacher-researchers across 8 U.S. universities and more than 400 computer and information science students, we embarked upon an Action Research investigation to gather insights into the pedagogical content knowledge (PCK) that teachers need to teach a particular inclusive design method called GenderMag. Analysis of the teachers' observations and experiences, the materials they used, direct observations of students' behaviors, and multiple data on the students' own reflections on their learning revealed 11 components of inclusive design PCK. These include strategies for anticipating and addressing resistance to the topic of inclusion, strategies for modeling and scaffolding perspective taking, and strategies for tailoring instruction to students' prior beliefs and biases.
Alannah Oleson, Christopher J. Mendez, Zoe Steine-Hanson, Claudia Hilderbrand, Christopher Perdriau, Margaret M. Burnett, Amy J. Ko
ICER6
2018 Open source barriers to entry, revisited: a sociotechnical perspective
abstract
Research has revealed that significant barriers exist when entering Open-Source Software (OSS) communities and that women disproportionately experience such barriers. However, this research has focused mainly on social/cultural factors, ignoring the environment itself --- the tools and infrastructure. To shed some light onto how tools and infrastructure might somehow factor into OSS barriers to entry, we conducted a field study with five teams of software professionals, who worked through five use-cases to analyze the tools and infrastructure used in their OSS projects. These software professionals found tool/infrastructure barriers in 7% to 71% of the use-case steps that they analyzed, most of which are tied to newcomer barriers that have been established in the literature. Further, over 80% of the barrier types they found include attributes that are biased against women.
Christopher J. Mendez, Hema Susmita Padala, Zoe Steine-Hanson, Claudia Hilderbrand, Amber Horvath, Charles Hill 0001, Logan Simpson, Nupoor Patil, Anita Sarma, Margaret M. Burnett
ICSE10
2018 Toward Foraging for Understanding of StarCraft Agents: An Empirical Study
abstract
Assessing and understanding intelligent agents is a difficult task for users that lack an AI background. A relatively new area, called "Explainable AI," is emerging to help address this problem, but little is known about how users would forage through information an explanation system might offer. To inform the development of Explainable AI systems, we conducted a formative study -- using the lens of Information Foraging Theory -- into how experienced users foraged in the domain of StarCraft to assess an agent. Our results showed that participants faced difficult foraging problems. These foraging problems caused participants to entirely miss events that were important to them, reluctantly choose to ignore actions they did not want to ignore, and bear high cognitive, navigation, and information costs to access the information they needed.
Sean Penney, Jonathan Dodge, Claudia Hilderbrand, Andrew Anderson 0002, Logan Simpson, Margaret M. Burnett
IUI6
2018 The GenderMag Recorder's Assistant
abstract
Building software systems is hard work, with challenges ranging from technical issues to usability issues. If the technical issues are not addressed, the software cannot work - but if the usability issues are not addressed, many potential users and customers are not even interested in whether it works. Further, usability must be inclusive: software needs to support diverse sorts of users. To help software professionals address gender-inclusive usability, we have created the GenderMag Recorder's Assistant tool. This Open Source tool is the first to semi-automate evaluating gender biases in software that is being designed, developed, or maintained. In this showpiece, we will demo the tool and encourage attendees to get involved in using it and improving upon it.
Christopher J. Mendez, Andrew Anderson 0002, Brijesh Bhuva, Margaret M. Burnett
VL/HCC4
2018 Semi-Automating (or not) a Socio-Technical Method for Socio-Technical Systems
abstract
How can we support software professionals who want to build human-adaptive sociotechnical systems? Building such systems requires skills some developers may lack, such as applying human-centric concepts to the software they develop and/or mentally modeling other people. Effective socio-technical methods exist to help, but most are manual and cognitively burdensome. In this paper, we investigate ways semi-automating a socio-technical method might help, using as our lens GenderMag, a method that requires people to mentally model people with genders different from their own. Toward this end, we created the GenderMag Recorder's Assistant, a semi-automated visual tool, and conducted a small field study and a 92-participant controlled study. Results of our investigation revealed ways the tool helped with cognitive load and ways it did not; unforeseen advantages of the tool in increasing participants' engagement with the method; and a few unforeseen advantages of the manual approach as well.
Christopher J. Mendez, Zoe Steine-Hanson, Alannah Oleson, Amber Horvath, Charles Hill 0001, Claudia Hilderbrand, Anita Sarma, Margaret M. Burnett
VL/HCC8
2017 Gender-Inclusiveness Personas vs. Stereotyping: Can We Have it Both Ways?
abstract
Personas often aim to improve product designers' ability to "see through the eyes of" target users through the empathy personas can inspire - but personas are also known to promote stereotyping. This tension can be particularly problematic when personas (who, of course as "people" have genders) are used to promote gender inclusiveness - because reinforcing stereotypical perceptions can run counter to gender inclusiveness. In this paper we explicitly investigate this tension through a new approach to personas: one that includes multiple photos (of males and females) for a single persona. We compared this approach to an identical persona with only one photo using a controlled laboratory study and an eye-tracking study. Our goal was to answer the following question: is it possible for personas to encourage product designers to engage with personas while at the same avoiding promoting gender stereotyping? Our results are encouraging about the use of personas with multiple pictures as a way to expand participants' consideration of multiple genders without reducing their engagement with the persona.
Charles Hill 0001, Maren Haag, Alannah Oleson, Christopher J. Mendez, Nicola Marsden, Anita Sarma, Margaret M. Burnett
CHI7
2017 PFIS-V: Modeling Foraging Behavior in the Presence of Variants
abstract
Foraging among similar variants of the same artifact is a common activity, but computational models of Information Foraging Theory (IFT) have not been developed to take such variants into account. Without being able to computationally predict people's foraging behavior with variants, our ability to harness the theory in practical ways--such as building and systematically assessing tools for people who forage different variants of an artifact--is limited. Therefore, in this paper, we introduce a new predictive model, PFIS-V, that builds upon PFIS3, the most recent of the PFIS family of modeling IFT in programming situations. Our empirical results show that PFIS-V is up to 25% more accurate than PFIS3 in predicting where a forager will navigate in a variationed information space.
Sruti Srinivasa Ragavan, Bhargav Pandya, David Piorkowski, Charles Hill 0001, Sandeep Kaur Kuttal, Anita Sarma, Margaret M. Burnett
CHI7
2017 Gender HCl and microsoft: Highlights from a longitudinal study
abstract
Research has emerged over the past decade showing gender biases in software. Although a few methods and prototype systems have emerged to help address this issue, none have been reported to have an impact on the people who actually build software. In this paper, we summarize a few highlights from a year-long field study investigating how Gender HCI methods to address gender biases in software can make impacts on a large software company.
Margaret M. Burnett, Robin Counts, Ronette Lawrence, Hannah Hanson
VL/HCC1
2017 Foraging goes mobile: Foraging while debugging on mobile devices
abstract
Although Information Foraging Theory (IFT) research for desktop environments has provided important insights into numerous information foraging tasks, we have been unable to locate IFT research for mobile environments. Despite the limits of mobile platforms, mobile apps are increasingly serving functions that were once exclusively the territory of desktops - and as the complexity of mobile apps increases, so does the need for foraging. In this paper we investigate, through a theory-based, dual replication study, whether and how foraging results from a desktop IDE generalize to a functionally similar mobile IDE. Our results show ways prior foraging research results from desktop IDEs generalize to mobile IDEs and ways they do not, and point to challenging open research questions for foraging on mobile environments.
David Piorkowski, Sean Penney, Austin Z. Henley, Marco Pistoia, Margaret M. Burnett, Omer Tripp, Pietro Ferrara 0001
VL/HCC5
2016 Finding Gender-Inclusiveness Software Issues with GenderMag: A Field Investigation
abstract
Gender inclusiveness in computing settings is receiving a lot of attention, but one potentially critical factor has mostly been overlooked -- software itself. To help close this gap, we recently created GenderMag, a systematic inspection method to enable software practitioners to evaluate their software for issues of gender-inclusiveness. In this paper, we present the first real-world investigation of software practitioners' ability to identify gender-inclusiveness issues in software they create/maintain using this method. Our investigation was a multiple-case field study of software teams at three major U.S. technology organizations. The results were that, using GenderMag to evaluate software, these software practitioners identified a surprisingly high number of gender-inclusiveness issues: 25% of the software features they evaluated had gender-inclusiveness issues.
Margaret M. Burnett, Anicia N. Peters, Charles Hill 0001, Noha Elarief
CHI1
2016 Programming, Problem Solving, and Self-Awareness: Effects of Explicit Guidance
abstract
More people are learning to code than ever, but most learning opportunities do not explicitly teach the problem solving skills necessary to succeed at open-ended programming problems. In this paper, we present a new approach to impart these skills, consisting of: 1) explicit instruction on programming problem solving, which frames coding as a process of translating mental representations of problems and solutions into source code, 2) a method of visualizing and monitoring progression through six problem solving stages, 3) explicit, on-demand prompts for learners to reflect on their strategies when seeking help from instructors, and 4) context-sensitive help embedded in a code editor that reinforces the problem solving instruction. We experimentally evaluated the effects of our intervention across two 2-week web development summer camps with 48 high school students, finding that the intervention increased productivity, independence, programming self-efficacy, metacognitive awareness, and growth mindset. We discuss the implications of these results on learning technologies and classroom instruction.
Dastyni Loksa, Amy J. Ko, Will Jernigan, Alannah Oleson, Christopher J. Mendez, Margaret M. Burnett
CHI6
2016 Foraging Among an Overabundance of Similar Variants
abstract
Foraging among too many variants of the same artifact can be problematic when many of these variants are similar. This situation, which is largely overlooked in the literature, is commonplace in several types of creative tasks, one of which is exploratory programming. In this paper, we investigate how novice programmers forage through similar variants. Based on our results, we propose a refinement to Information Foraging Theory (IFT) to include constructs about variation foraging behavior, and propose refinements to computational models of IFT to better account for foraging among variants.
Sruti Srinivasa Ragavan, Sandeep Kaur Kuttal, Charles Hill 0001, Anita Sarma, David Piorkowski, Margaret M. Burnett
CHI6
2016 "Womenomics" and gender-inclusive software: what software engineers need to know (invited talk)
abstract
This short paper is a summary of my keynote at FSE’16, with accompanying references for follow-up.
Margaret M. Burnett
SIGSOFT FSE1
2016 Foraging and navigations, fundamentally: developers' predictions of value and cost
abstract
Empirical studies have revealed that software developers spend 35%–50% of their time navigating through source code during development activities, yet fundamental questions remain: Are these percentages too high, or simply inherent in the nature of software development? Are there factors that somehow determine a lower bound on how effectively developers can navigate a given information space? Answering questions like these requires a theory that captures the core of developers' navigation decisions. Therefore, we use the central proposition of Information Foraging Theory to investigate developers' ability to predict the value and cost of their navigation decisions. Our results showed that over 50% of developers' navigation choices produced less value than they had predicted and nearly 40% cost more than they had predicted. We used those results to guide a literature analysis, to investigate the extent to which these challenges are met by current research efforts, revealing a new area of inquiry with a rich and crosscutting set of research challenges and open problems.
David Piorkowski, Austin Z. Henley, Tahmid Nabi, Scott D. Fleming, Christopher Scaffidi, Margaret M. Burnett
SIGSOFT FSE6
2016 GenderMag experiences in the field: The whole, the parts, and the workload
abstract
Recent research has reported numerous studies bringing into question the gender inclusiveness of many kinds of software. Inclusiveness of software (gender or otherwise) matters because supporting diversity matters - it is well-known that the more diverse a group of problem-solvers, the higher the quality of the solution. To help software creators identify features within their software that are not gender-inclusive, we recently created a method known as GenderMag. In this paper, we investigate the experience of teams of software professionals using GenderMag to find problems with software they are building. Our results show a high engagement with GenderMag personas - more than twice that of other personas research - and a very high degree of accuracy (93%) most of the time. Finally, our results pinpointed situations that we term “detours” that were especially prone to errors, with teams 6 times more likely to make errors in detours than they did otherwise.
Charles Hill 0001, Shannon Ernst, Alannah Oleson, Amber Horvath, Margaret M. Burnett
VL/HCC5
2016 Trials and tribulations of developers of intelligent systems: A field study
abstract
Intelligent systems are gaining in popularity and receiving increased media attention, but little is known about how people actually go about developing them. In this paper, we attempt to fill this gap through a set of field interviews that investigate how people develop intelligent systems that incorporate machine learning algorithms. The developers we interviewed were experienced at working with machine learning algorithms and dealing with the large amounts of data needed to develop intelligent systems. Despite their level of experience, we learned that they struggle to establish a repeatable process. They described problems with each step of the processes they perform, as well as cross-cutting issues that pervade multiple steps of their processes. The unique difficulties that developers like these face seem to point to a need for software engineering advances that address such machine learning systems, and we conclude by discussing this need and some of its implications.
Charles Hill 0001, Rachel K. E. Bellamy, Thomas Erickson, Margaret M. Burnett
VL/HCC4
2016 Putting information foraging theory to work: Community-based design patterns for programming tools
abstract
The design of programming tools is slow and costly. To ease this process, we developed a design pattern catalog aimed at providing guidance for tool designers. This catalog is grounded in Information Foraging Theory (IFT), which empirical studies have shown to be useful for understanding how developers look for information during development tasks. New design patterns, authored by members of the research community for the catalog, concretely explain how to apply IFT in tool design. In our evaluation, qualitative analyses revealed the community-written design patterns compared well in quality to patterns that we had ourselves published in a smaller, peer-reviewed catalog.
Tahmid Nabi, Kyle M. D. Sweeney, Sam Lichlyter, David Piorkowski, Christopher Scaffidi, Margaret M. Burnett, Scott D. Fleming
VL/HCC6
2016 GenderMag: A Method for Evaluating Software's Gender Inclusiveness
abstract
In recent years, research into gender differences has established that individual differences in how people problem-solve often cluster by gender. Research also shows that these differences have direct implications for software that aims to support users' problem-solving activities, and that much of this software is more supportive of problem-solving processes favored (statistically) more by males than by females. However, there is almost no work considering how software practitioners—such as User Experience (UX) professionals or software developers—can find gender-inclusiveness issues like these in their software. To address this gap, we devised the GenderMag method for evaluating problem-solving software from a gender-inclusiveness perspective. The method includes a set of faceted personas that bring five facets of gender difference research to life, and embeds use of the personas into a concrete process through a gender-specialized Cognitive Walkthrough. Our empirical results show that a variety of practitioners who design software—without needing any background in gender research—were able to use the GenderMag method to find gender-inclusiveness issues in problem-solving software. Our results also show that the issues the practitioners found were real and fixable. This work is the first systematic method to find gender-inclusiveness issues in software, so that practitioners can design and produce problem-solving software that is more usable by everyone.
Margaret M. Burnett, Simone Stumpf, Stephann Makri, Laura Beckwith, Irwin Kwan, Anicia N. Peters, Will Jernigan
Interact. Comput.1
2015 To fix or to learn? How production bias affects developers' information foraging during debugging
abstract
Developers performing maintenance activities must balance their efforts to learn the code vs. their efforts to actually change it. This balancing act is consistent with the “production bias” that, according to Carroll's minimalist learning theory, generally affects software users during everyday tasks. This suggests that developers' focus on efficiency should have marked effects on how they forage for the information they think they need to fix bugs. To investigate how developers balance fixing versus learning during debugging, we conducted the first empirical investigation of the interplay between production bias and information foraging. Our theory-based study involved 11 participants: half tasked with fixing a bug, and half tasked with learning enough to help someone else fix it. Despite the subtlety of difference between their tasks, participants foraged remarkably differently-making foraging decisions from different types of “patches,” with different types of information, and succeeding with different foraging tactics.
David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Margaret M. Burnett, Irwin Kwan, Austin Z. Henley, Charles Hill 0001, Amber Horvath
ICSME4
2015 Principles of Explanatory Debugging to Personalize Interactive Machine Learning
abstract
How can end users efficiently influence the predictions that machine learning systems make on their behalf? This paper presents Explanatory Debugging, an approach in which the system explains to users how it made each of its predictions, and the user then explains any necessary corrections back to the learning system. We present the principles underlying this approach and a prototype instantiating it. An empirical evaluation shows that Explanatory Debugging increased participants' understanding of the learning system by 52% and allowed participants to correct its mistakes up to twice as efficiently as participants using a traditional learning system.
Todd Kulesza, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf
IUI2
2015 A principled evaluation for a principled idea garden
abstract
Many systems are designed to help novices who want to learn programming, but few support those who are not interested in learning (more) programming. This paper targets the subset of end-user programmers (EUPs) in this category. We present a set of principles on how to help EUPs like this learn just a little when they need to overcome a barrier. We then instantiate the principles in a prototype and empirically investigate the principles in two studies: a formative think-aloud study and a pair of summer camps attended by 42 teens. Among the surprising results were the complementary roles of implicitly actionable hints versus explicitly actionable hints, and the importance of both context-free and context-sensitive availability. Under these principles, the camp participants required significantly less in-person help than in a previous camp to learn the same amount of material in the same amount of time.
Will Jernigan, Amber Horvath, Michael Jongseon Lee, Margaret M. Burnett, Taylor Cuilty, Sandeep Kaur Kuttal, Anicia N. Peters, Irwin Kwan, Faezeh Bahmani, Amy J. Ko
VL/HCC4
2015 A practical guide to controlled experiments of software engineering tools with human participants
Amy J. Ko, Thomas D. LaToza, Margaret M. Burnett
Empir. Softw. Eng.3
2015 Idea Garden: Situated Support for Problem Solving by End-User Programmers
abstract
Although there have been many advances in end-user programming environments, recent empirical studies report that programming still remains difficult for end-users. We hypothesize that one reason may be lack of effective support for helping end-user programmers problem-solve their own way around barriers they encounter. Therefore, in this paper, we describe the Idea Garden, a concept designed to help end-user programmers generate new ideas and problem-solve when they run into barriers. The Idea Garden has its roots in Minimalist Learning Theory and problem-solving theories. Our proof-of-concept prototype of the Idea Garden concept in the CoScripter end-user programming environment currently targets three barriers reported in end-user programming literature. It does so using an integrated, just-in-time combination of scaffolding for problem-solving strategies, for design patterns and for programming concepts. Our empirical results showed that this approach helped end-user programmers overcome all three types of barriers that our prototype targeted.
Jill Cao, Scott D. Fleming, Margaret M. Burnett, Christopher Scaffidi
Interact. Comput.3
2014 Principles of a debugging-first puzzle game for computing education
abstract
Although there are many systems designed to engage people in programming, few explicitly teach the subject, expecting learners to acquire the necessary skills on their own as they create programs from scratch. We present a principled approach to teach programming using a debugging game called Gidget, which was created using a unique set of seven design principles. A total of 44 teens played it via a lab study and two summer camps. Principle by principle, the results revealed strengths, problems, and open questions for the seven principles. Taken together, the results were very encouraging: learners were able to program with conditionals, loops, and other programming concepts after using the game for just 5 hours.
Michael Jongseon Lee, Faezeh Bahmani, Irwin Kwan, Jilian LaFerte, Polina Charters, Amber Horvath, Fanny Luor, Jill Cao, Catherine Law, Michael Beswetherick, Sheridan Long, Margaret M. Burnett, Amy J. Ko
VL/HCC12
2014 You Are the Only Possible Oracle: Effective Test Selection for End Users of Interactive Machine Learning Systems
abstract
How do you test a program when only a single user, with no expertise in software testing, is able to determine if the program is performing correctly? Such programs are common today in the form of machine-learned classifiers. We consider the problem of testing this common kind of machine-generated program when the only oracle is an end user: e.g., only you can determine if your email is properly filed. We present test selection methods that provide very good failure rates even for small test suites, and show that these methods work in both large-scale random experiments using a “gold standard” and in studies with real users. Our methods are inexpensive and largely algorithm-independent. Key to our methods is an exploitation of properties of classifiers that is not possible in traditional software testing. Our results suggest that it is plausible for time-pressured end users to interactively detect failures-even very hard-to-find failures-without wading through a large number of successful (and thus less useful) tests. We additionally show that some methods are able to find the arguably most difficult-to-detect faults of classifiers: cases where machine learning algorithms have high confidence in an incorrect result.
Alex Groce, Todd Kulesza, Chaoqiang Zhang, Shalini Shamasunder, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf, Shubhomoy Das, Amber Shinsel, Forrest Bice, Kevin McIntosh
IEEE Trans. Software Eng.5
2013 The whats and hows of programmers' foraging diets
abstract
One of the least studied areas of Information Foraging Theory is diet: the information foragers choose to seek. For example, do foragers choose solely based on cost, or do they stubbornly pursue certain diets regardless of cost? Do their debugging strategies vary with their diets? To investigate "what" and "how" questions like these for the domain of software debugging, we qualitatively analyzed 9 professional developers' foraging goals, goal patterns, and strategies. Participants spent 50% of their time foraging. Of their foraging, 58% fell into distinct dietary patterns - mostly in patterns not previously discussed in the literature. In general, programmers' foraging strategies leaned more heavily toward enrichment than we expected, but different strategies aligned with different goal types. These and our other findings help fill the gap as to what programmers' dietary goals are and how their strategies relate to those goals.
David Piorkowski, Scott D. Fleming, Irwin Kwan, Margaret M. Burnett, Christopher Scaffidi, Rachel K. E. Bellamy, Joshua Jordahl
CHI4
2013 End-user programmers in trouble: Can the Idea Garden help them to help themselves?
abstract
End-user programmers often get stuck because they do not know how to overcome their barriers. We have previously presented an approach called the Idea Garden, which makes minimalist, on-demand problem-solving support available to end-user programmers in trouble. Its goal is to encourage end users to help themselves learn how to overcome programming difficulties as they encounter them. In this paper, we investigate whether the Idea Garden approach helps end-user programmers problem-solve their programs on their own. We ran a statistical experiment with 123 end-user programmers. The experiment's results showed that, even when the Idea Garden was no longer available, participants with little knowledge of programming who previously used the Idea Garden were able to produce higher-quality programs than those who had not used the Idea Garden.
Jill Cao, Irwin Kwan, Faezeh Bahmani, Margaret M. Burnett, Scott D. Fleming, Joshua Jordahl, Amber Horvath, Sherry Yang 0002
VL/HCC4
2013 Too much, too little, or just right? Ways explanations impact end users' mental models
abstract
Research is emerging on how end users can correct mistakes their intelligent agents make, but before users can correctly “debug” an intelligent agent, they need some degree of understanding of how it works. In this paper we consider ways intelligent agents should explain themselves to end users, especially focusing on how the soundness and completeness of the explanations impacts the fidelity of end users' mental models. Our findings suggest that completeness is more important than soundness: increasing completeness via certain information types helped participants' mental models and, surprisingly, their perception of the cost/benefit tradeoff of attending to the explanations. We also found that oversimplification, as per many commercial agents, can be a problem: when soundness was very low, participants experienced more mental demand and lost trust in the explanations, thereby reducing the likelihood that users will pay attention to such explanations at all.
Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Sherry Yang 0002, Irwin Kwan, Weng-Keen Wong
VL/HCC3
2013 End-user feature labeling: Supervised and semi-supervised approaches based on locally-weighted logistic regression
Shubhomoy Das, Travis Moore, Weng-Keen Wong, Simone Stumpf, Ian Oberst, Kevin McIntosh, Margaret M. Burnett
Artif. Intell.7
2013 An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks
abstract
Theories of human behavior are an important but largely untapped resource for software engineering research. They facilitate understanding of human developers’ needs and activities, and thus can serve as a valuable resource to researchers designing software engineering tools. Furthermore, theories abstract beyond specific methods and tools to fundamental principles that can be applied to new situations. Toward filling this gap, we investigate the applicability and utility of Information Foraging Theory (IFT) for understanding information-intensive software engineering tasks, drawing upon literature in three areas: debugging, refactoring, and reuse. In particular, we focus on software engineering tools that aim to support information-intensive activities, that is, activities in which developers spend time seeking information. Regarding applicability, we consider whether and how the mathematical equations within IFT can be used to explain why certain existing tools have proven empirically successful at helping software engineers. Regarding utility, we applied an IFT perspective to identify recurring design patterns in these successful tools, and consider what opportunities for future research are revealed by our IFT perspective.
Scott D. Fleming, Christopher Scaffidi, David Piorkowski, Margaret M. Burnett, Rachel K. E. Bellamy, Joseph Lawrance, Irwin Kwan
ACM Trans. Softw. Eng. Methodol.4
2013 How Programmers Debug, Revisited: An Information Foraging Theory Perspective
abstract
Many theories of human debugging rely on complex mental constructs that offer little practical advice to builders of software engineering tools. Although hypotheses are important in debugging, a theory of navigation adds more practical value to our understanding of how programmers debug. Therefore, in this paper, we reconsider how people go about debugging in large collections of source code using a modern programming environment. We present an information foraging theory of debugging that treats programmer navigation during debugging as being analogous to a predator following scent to find prey in the wild. The theory proposes that constructs of scent and topology provide enough information to describe and predict programmer navigation during debugging, without reference to mental states such as hypotheses. We investigate the scope of our theory through an empirical study of 10 professional programmers debugging a real-world open source program. We found that the programmers' verbalizations far more often concerned scent-following than hypotheses. To evaluate the predictiveness of our theory, we created an executable model that predicted programmer navigation behavior more accurately than comparable models that did not consider information scent. Finally, we discuss the implications of our results for enhancing software engineering tools.
Joseph Lawrance, Christopher Bogart, Margaret M. Burnett, Rachel K. E. Bellamy, Kyle Rector, Scott D. Fleming
IEEE Trans. Software Eng.3
2012 Designing a debugging interaction language for cognitive modelers: an initial case study in natural programming plus
abstract
In this paper, we investigate how a debugging environment should support a population doing work at the core of HCI research: cognitive modelers. In conducting this investigation, we extended the Natural Programming methodology (a user-centered design method for HCI researchers of programming environments), to add an explicit method for mapping the outcomes of NP's empirical investigations to a language design. This provided us with a concrete way to make the design leap from empirical assessment of users' needs to a language. The contributions of our work are therefore: (1) empirical evidence about the content and sequence of cognitive modelers' information needs when debugging, (2) a new, empirically derived, design specification for a debugging interaction language for cognitive modelers, and (3) an initial case study of our "Natural Programming Plus" methodology.
Christopher Bogart, Margaret M. Burnett, Scott Douglass, Hannah Adams, Rachel White
CHI2
2012 Tell me more?: the effects of mental model soundness on personalizing an intelligent agent
abstract
What does a user need to know to productively work with an intelligent agent? Intelligent agents and recommender systems are gaining widespread use, potentially creating a need for end users to understand how these systems operate in order to fix their agent's personalized behavior. This paper explores the effects of mental model soundness on such personalization by providing structural knowledge of a music recommender system in an empirical study. Our findings show that participants were able to quickly build sound mental models of the recommender system's reasoning, and that participants who most improved their mental models during the study were significantly more likely to make the recommender operate to their satisfaction. These results suggest that by helping end users understand a system's reasoning, intelligent agents may elicit more and better feedback, thus more closely aligning their output with each user's intentions.
Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Irwin Kwan
CHI3
2012 Reactive information foraging: an empirical investigation of theory-based recommender systems for programmers
abstract
Information Foraging Theory (IFT) has established itself as an important theory to explain how people seek information, but most work has focused more on the theory itself than on how best to apply it. In this paper, we investigate how to apply a reactive variant of IFT (Reactive IFT) to design IFT-based tools, with a special focus on such tools for ill-structured problems. Toward this end, we designed and implemented a variety of recommender algorithms to empirically investigate how to help people with the ill-structured problem of finding where to look for information while debugging source code. We varied the algorithms based on scent type supported (words alone vs. words + code structure), and based on use of foraging momentum to estimate rapidity of foragers' goal changes. Our empirical results showed that (1) using both words and code structure significantly improved the ability of the algorithms to recommend where software developers should look for information; (2) participants used recommendations to discover new places in the code and also as shortcuts to navigate to known places; and (3) low-momentum recommendations were significantly more useful than high-momentum recommendations, suggesting rapid and numerous goal changes in this type of setting. Overall, our contributions include two new recommendation algorithms, empirical evidence about when and why participants found IFT-based recommendations useful, and implications for the design of tools based on Reactive IFT.
David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Bonnie E. John, Rachel K. E. Bellamy, Calvin Swart
CHI5
2012 Towards recognizing "cool": can end users help computer vision recognize subjective attributes of objects in images?
abstract
Recent computer vision approaches are aimed at richer image interpretations that extend the standard recognition of objects in images (e.g., cars) to also recognize object attributes (e.g., cylindrical, has-stripes, wet). However, the more idiosyncratic and abstract the notion of an object attribute (e.g., cool car), the more challenging the task of attribute recognition. This paper considers whether end users can help vision algorithms recognize highly idiosyncratic attributes, referred to here as subjective attributes. We empirically investigated how end users recognized three subjective attributes of carscool, cute, and classic. Our results suggest the feasibility of vision algorithms recognizing subjective attributes of objects, but an interactive approach beyond standard supervised learning from labeled training examples is needed.
William Curran, Travis Moore, Todd Kulesza, Weng-Keen Wong, Sinisa Todorovic, Simone Stumpf, Rachel White, Margaret M. Burnett
IUI8
2012 From barriers to learning in the idea garden: An empirical study
abstract
How can end-user programming environments better help their users overcome programming barriers? We have been investigating an approach called Idea Gardening, which addresses this problem by helping end users to help themselves overcome barriers in the context of “doing”. In this paper, we report on a qualitative empirical study of how effectively an Idea Garden prototype helped end users overcome programming barriers in the CoScripter environment, and the extent to which participants learned after interacting with our features. Our results showed that 9 out of 10 participants who encountered barriers and then used the Idea Garden, overcame their barriers. Further, all 9 went on to demonstrate evidence of having learned the programming concepts, patterns, and strategies relevant to overcoming these barriers.
Jill Cao, Irwin Kwan, Rachel White, Scott D. Fleming, Margaret M. Burnett, Christopher Scaffidi
VL/HCC5
2012 End-user debugging strategies: A sensemaking perspective
abstract
Despite decades of research into how professional programmers debug, only recently has work emerged about how end-user programmers attempt to debug programs. Without this knowledge, we cannot build tools to adequately support their needs. This article reports the results of a detailed qualitative empirical study of end-user programmers' sensemaking about a spreadsheet's correctness. Using our study's data, we derived a sensemaking model for end-user debugging and categorized participants' activities and verbalizations according to this model, allowing us to investigate how participants went about debugging. Among the results are identification of the prevalence of information foraging during end-user debugging, two successful strategies for traversing the sensemaking model, potential ties to gender differences in the literature, sensemaking sequences leading to debugging progress, and sequences tied with troublesome points in the debugging process. The results also reveal new implications for the design of spreadsheet tools to support end-user programmers' sensemaking during debugging.
Valentina Grigoreanu, Margaret M. Burnett, Susan Wiedenbeck, Jill Cao, Kyle Rector, Irwin Kwan
ACM Trans. Comput. Hum. Interact.2
2011 End-User Feature Labeling via Locally Weighted Logistic Regression
abstract
Applications that adapt to a particular end user often make inaccurate predictions during the early stages when training data is limited. Although an end user can improve the learning algorithm by labeling more training data, this process is time consuming and too ad hoc to target a particular area of inaccuracy. To solve this problem, we propose a new learning algorithm based on Locally Weighted Logistic Regression for feature labeling by end users, enabling them to point out which features are important for a class, rather than provide new training instances. In our user study, the first allowing ordinary end users to freely choose features to label directly from text documents, our algorithm was more effective than others at leveraging end users’ feature labels to improve the learning algorithm. Our results strongly suggest that allowing users to freely choose features to label is a promising method for allowing end users to improve learning algorithms effectively.
Weng-Keen Wong, Ian Oberst, Shubhomoy Das, Travis Moore, Simone Stumpf, Kevin McIntosh, Margaret M. Burnett
AAAI7
2011 End-user feature labeling: a locally-weighted regression approach
abstract
When intelligent interfaces, such as intelligent desktop assistants, email classifiers, and recommender systems, customize themselves to a particular end user, such customizations can decrease productivity and increase frustration due to inaccurate predictions - especially in early stages, when training data is limited. The end user can improve the learning algorithm by tediously labeling a substantial amount of additional training data, but this takes time and is too ad hoc to target a particular area of inaccuracy. To solve this problem, we propose a new learning algorithm based on locally weighted regression for feature labeling by end users, enabling them to point out which features are important for a class, rather than provide new training instances. In our user study, the first allowing ordinary end users to freely choose features to label directly from text documents, our algorithm was both more effective than others at leveraging end users' feature labels to improve the learning algorithm, and more robust to real users' noisy feature labels. These results strongly suggest that allowing users to freely choose features to label is a promising method for allowing end users to improve learning algorithms effectively.
Weng-Keen Wong, Ian Oberst, Shubhomoy Das, Travis Moore, Simone Stumpf, Kevin McIntosh, Margaret M. Burnett
IUI7
2011 An exploration of design opportunities for "gardening" end-user programmers' ideas
abstract
Despite recent advances in supporting end-user programmers, empirical studies continue to report barriers that end users experience in problem solving with programming environments. We hypothesize that an important barrier that still needs to be overcome is the lack of support for nurturing end-user programmers' ideas on how a program should be written or on how to solve programming difficulties. Therefore, in this paper, we present a qualitative empirical investigation and triangulate the results with theories from problem solving and creativity. Moreover, we explore design opportunities and a design space for “idea gardening”, a new approach to nurturing end-user programmers' ideas and to helping them gradually gain expertise as they overcome barriers. Our results suggest that nurturing end-user programmers' ideas is a fertile area for research with an interesting, multidimensional design space.
Jill Cao, Scott D. Fleming, Margaret M. Burnett
VL/HCC3
2011 Modeling programmer navigation: A head-to-head empirical evaluation of predictive models
abstract
Software developers frequently need to perform code maintenance tasks, but doing so requires time-consuming navigation through code. A variety of tools are aimed at easing this navigation by using models to identify places in the code that a developer might want to visit, and then providing shortcuts so that the developer can quickly navigate to those locations. To date, however, only a few of these models have been compared head-to-head to assess their predictive accuracy. In particular, we do not know which models are most accurate overall, which are accurate only in certain circumstances, and whether combining models could enhance accuracy. Therefore, we have conducted an empirical study to evaluate the accuracy of a broad range of models for predicting many different kinds of code navigations in sample maintenance tasks. Overall, we found that models tended to perform best if they took into account how recently a developer has viewed pieces of the code, and if models took into account the spatial proximity of methods within the code. We also found that the accuracy of single-factor models can be improved by combining factors, using a spreading-activation based approach, to produce multi-factor models. Based on these results, we offer concrete guidance about how these models could be used to provide enhanced software development tools that ease the difficulty of navigating through code.
David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Liza John, Christopher Bogart, Bonnie E. John, Margaret M. Burnett, Rachel K. E. Bellamy
VL/HCC7
2011 Mini-crowdsourcing end-user assessment of intelligent assistants: A cost-benefit study
abstract
Intelligent assistants sometimes handle tasks too important to be trusted implicitly. End users can establish trust via systematic assessment, but such assessment is costly. This paper investigates whether, when, and how bringing a small crowd of end users to bear on the assessment of an intelligent assistant is useful from a cost/benefit perspective. Our results show that a mini-crowd of testers supplied many more benefits than the obvious decrease in workload, but these benefits did not scale linearly as mini-crowd size increased - there was a point of diminishing returns where the cost-benefit ratio became less attractive.
Amber Shinsel, Todd Kulesza, Margaret M. Burnett, William Curran, Alex Groce, Simone Stumpf, Weng-Keen Wong
VL/HCC3
2011 Gender pluralism in problem-solving software
abstract
Although there has been significant research into gender regarding educational and workplace practices, there has been little awareness of gender differences as they pertain to software tools, such as spreadsheet applications, that try to support end users in problem-solving tasks. Although such software tools are intended to be gender agnostic, we believe that closer examination of this premise is warranted. Therefore, in this paper, we report an end-to-end investigation into gender differences with spreadsheet software. Our results showed gender differences in feature usage, feature-related confidence, and tinkering (playful exploration) with features. Then, drawing implications from these results, we designed and implemented features for our spreadsheet prototype that took the gender differences into account. The results of an evaluation on this prototype showed improvements for both males and females, and also decreased gender differences in some outcome measures, such as confidence. These results are encouraging, but also open new questions for investigation. We also discuss how our results compare to generalization studies performed with a variety of other software platforms and populations.
Margaret M. Burnett, Laura Beckwith, Susan Wiedenbeck, Scott D. Fleming, Jill Cao, Thomas H. Park, Valentina Grigoreanu, Kyle Rector
Interact. Comput.1
2011 Why-oriented end-user debugging of naive Bayes text classification
abstract
Machine learning techniques are increasingly used in intelligent assistants , that is, software targeted at and continuously adapting to assist end users with email, shopping, and other tasks. Examples include desktop SPAM filters, recommender systems, and handwriting recognition. Fixing such intelligent assistants when they learn incorrect behavior, however, has received only limited attention. To directly support end-user “debugging” of assistant behaviors learned via statistical machine learning, we present a Why-oriented approach which allows users to ask questions about how the assistant made its predictions, provides answers to these “why” questions, and allows users to interactively change these answers to debug the assistant's current and future predictions. To understand the strengths and weaknesses of this approach, we then conducted an exploratory study to investigate barriers that participants could encounter when debugging an intelligent assistant using our approach, and the information those participants requested to overcome these barriers. To help ensure the inclusiveness of our approach, we also explored how gender differences played a role in understanding barriers and information needs. We then used these results to consider opportunities for Why-oriented approaches to address user barriers and information needs.
Todd Kulesza, Simone Stumpf, Weng-Keen Wong, Margaret M. Burnett, Stephen Perona, Amy J. Ko, Ian Oberst
ACM Trans. Interact. Intell. Syst.4
2010 End-user mashup programming: through the design lens
abstract
Programming has recently become more common among ordinary end users of computer systems. We believe that these end-user programmers are not just coders but also designers, in that they interlace making design decisions with coding rather than treating them as two separate phases. To better understand and provide support for the programming and design needs of end users, we propose a design theory-based approach to look at end-user programming. Toward this end, we conducted a think-aloud study with ten end users creating a web mashup. By analyzing users' verbal and behavioral data using Schon's reflection-in-action design model and the notion of ideations from creativity literature, we discovered insights into end-user programmers' problem-solving attempts, successes, and obstacles, with accompanying implications for the design of end-user programming environments for mashups. The contribution of our work is three-fold: 1) the methodology of using a design lens to view programming, 2) evidence, through insights gained, of the usefulness of this approach, and 3) the implications themselves.
Jill Cao, Yann Riche, Susan Wiedenbeck, Margaret M. Burnett, Valentina Grigoreanu
CHI4
2010 A strategy-centric approach to the design of end-user debugging tools
abstract
End-user programmers' code is notoriously buggy. This problem is amplified by the increasing complexity of end users' programs. To help end users catch errors early and reliably, we employ a novel approach for the design of end-user debugging tools: a focus on supporting end users' effective debugging strategies. This paper makes two contributions. We first demonstrate the potential of a strategy-centric approach to tool design by presenting StratCel, an add-in for Excel. Second, we show the benefits of this design approach: participants using StratCel found twice as many bugs as participants using standard Excel, they fixed four times as many bugs, and all this in only a small fraction of the time. Other contributions included: a boost in novices' debugging performance near experienced participants' improved levels, validated design guidelines, a discussion of the generalizability of this approach, and several opportunities for future research.
Valentina Grigoreanu, Margaret M. Burnett, George G. Robertson
CHI2
2010 Reactive information foraging for evolving goals
abstract
Information foraging models have predicted the navigation paths of people browsing the web and (more recently) of programmers while debugging, but these models do not explicitly model users' goals evolving over time. We present a new information foraging model called PFIS2 that does model information seeking with potentially evolving goals. We then evaluated variants of this model in a field study that analyzed programmers' daily navigations over a seven-month period. Our results were that PFIS2 predicted users' navigation remarkably well, even though the goals of navigation, and even the information landscape itself, were changing markedly during the pursuit of information.
Joseph Lawrance, Margaret M. Burnett, Rachel K. E. Bellamy, Christopher Bogart, Calvin Swart
CHI2
2010 Gender differences and programming environments: across programming populations
abstract
Although there has been significant research into gender regarding educational and workplace practices, there has been little investigation of gender differences pertaining to problem solving with programming tools and environments. As a result, there is little evidence as to what role gender plays in programming tools---and what little evidence there is has involved mainly novice and end-user programmers in academic studies. This paper therefore investigates how widespread such phenomena are in industrial programming situations, considering three disparate programming populations involving almost 3000 people and three different programming platforms in industry. To accomplish this, we analyzed four industry "legacy" studies from a gender perspective, triangulating results against each other and against a new fifth study, also in industry. We investigated gender differences in software feature usage and in tinkering/exploring software features. Furthermore, we examined how such differences tied to confidence. Our results showed significant gender differences in all three factors---across all populations and platforms.
Margaret M. Burnett, Scott D. Fleming, Shamsi T. Iqbal, Gina Venolia, Vidya Rajaram, Valentina Grigoreanu, Mary Czerwinski
ESEM1
2010 Does My Model Work? Evaluation Abstractions of Cognitive Modelers
abstract
Are the abstractions that scientific modelers use to build their models in a modeling language the same abstractions they use to evaluate the correctness of their models? The extent to which such differences exist seems likely to correspond to additional effort of modelers in determining whether their models work as intended. In this paper, we therefore investigate the distinction between "programming abstractions" and "evaluation abstractions". As the basis of our investigation, we conducted a case study on cognitive modeling. We report modelers' evaluation abstractions, and the lengths they went to in evaluating their models. From these results, we derive design implications for several categories of persistent, first-class evaluation abstractions in future debugging tools for modelers.
Christopher Bogart, Margaret M. Burnett, Scott Douglass, David Piorkowski, Amber Shinsel
VL/HCC2
2010 A Debugging Perspective on End-User Mashup Programming
abstract
In recent years, systems have emerged that enable end users to “mash” together existing web services to build new web sites. However, little is known about how well end users succeed at building such mashups, or what they do if they do not succeed at their first attempt. To help fill this gap, we took a fresh look, from a debugging perspective, at the approaches of end users as they attempted to create mashups. Our results reveal the end users' debugging strategies and strategy barriers, the gender differences between the debugging strategies males and females followed and the features they used, and finally how their debugging successes and difficulties interacted with their design behaviors.
Jill Cao, Kyle Rector, Thomas H. Park, Scott D. Fleming, Margaret M. Burnett, Susan Wiedenbeck
VL/HCC5
2010 Explanatory Debugging: Supporting End-User Debugging of Machine-Learned Programs
abstract
Many machine-learning algorithms learn rules of behavior from individual end users, such as task-oriented desktop organizers and handwriting recognizers. These rules form a “program” that tells the computer what to do when future inputs arrive. Little research has explored how an end user can debug these programs when they make mistakes. We present our progress toward enabling end users to debug these learned programs via a Natural Programming methodology. We began with a formative study exploring how users reason about and correct a text-classification program. From the results, we derived and prototyped a concept based on “explanatory debugging”, then empirically evaluated it. Our results contribute methods for exposing a learned program's logic to end users and for eliciting user corrections to improve the program's predictions.
Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Weng-Keen Wong, Yann Riche, Travis Moore, Ian Oberst, Amber Shinsel, Kevin McIntosh
VL/HCC3
2010 Mining problem-solving strategies from HCI data
abstract
Can we learn about users' problem-solving strategies by observing their actions? This article introduces a data mining system that extracts complex behavioral patterns from logged user actions to discover users' high-level strategies. Our application domain is an HCI study aimed at revealing users' strategies in an end-user debugging task and understanding how the strategies relate to gender and to success. We cast this problem as a sequential pattern discovery problem, where user strategies are manifested as sequential behavior patterns. Problematically, we found that the patterns discovered by standard data mining algorithms were difficult to interpret and provided limited information about high-level strategies. To help interpret the patterns as strategies, we examined multiple ways of clustering the patterns into meaningful groups. This collectively led to interesting findings about users' behavior in terms of both gender differences and debugging success. These common behavioral patterns were novel HCI findings about differences in males' and females' behavior with software, and were verified by a parallel study with an independent data set on strategies. As a research endeavor into the interpretability issues faced by data mining techniques, our work also highlights important research directions for making data mining more accessible to non-data-mining experts.
Xiaoli Z. Fern, Chaitanya Komireddy, Valentina Grigoreanu, Margaret M. Burnett
ACM Trans. Comput. Hum. Interact.4
2009 Fixing the program my computer learned: barriers for end users, challenges for the machine
abstract
The results of a machine learning from user behavior can be thought of as a program, and like all programs, it may need to be debugged. Providing ways for the user to debug it matters, because without the ability to fix errors users may find that the learned program's errors are too damaging for them to be able to trust such programs. We present a new approach to enable end users to debug a learned program. We then use an early prototype of our new approach to conduct a formative study to determine where and when debugging issues arise, both in general and also separately for males and females. The results suggest opportunities to make machine-learned programs more effective tools.
Todd Kulesza, Weng-Keen Wong, Simone Stumpf, Stephen Perona, Rachel White, Margaret M. Burnett, Ian Oberst, Amy J. Ko
IUI6
2009 Predicting reuse of end-user web macro scripts
abstract
Repositories of code written by end-user programmers are beginning to emerge, but when a piece of code is new or nobody has yet reused it, then current repositories provide users with no information about whether that code might be appropriate for reuse. Addressing this problem requires predicting reusability based on information that exists when a script is created. To provide such a model for web macro scripts, we identified script traits that might plausibly predict reuse, then used IBM CoScripter repository logs to statistically test how well each corresponded to reuse. We then built a machine learning model that combines the useful traits and evaluated how well it can predict four different types of reuse that we saw in the repository logs. Our model was able to predict reuse from a surprisingly small set of traits. It is simple enough to be explained in only 6-11 rules, making it potentially viable for integration in repository search engines for end-user programmers.
Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Allen Cypher, Brad A. Myers, Mary Shaw
VL/HCC3
2009 Interacting meaningfully with machine learning systems: Three experiments
Simone Stumpf, Vidya Rajaram, Lida Li, Weng-Keen Wong, Margaret M. Burnett, Thomas G. Dietterich, Erin Sullivan, Jon Herlocker
Int. J. Hum. Comput. Stud.5
2008 Using information scent to model the dynamic foraging behavior of programmers in maintenance tasks
abstract
In recent years, the software engineering community has begun to study program navigation and tools to support it. Some of these navigation tools are very useful, but they lack a theoretical basis that could reduce the need for ad hoc tool building approaches by explaining what is fundamentally necessary in such tools. In this paper, we present PFIS (Programmer Flow by Information Scent), a model and algorithm of programmer navigation during software maintenance. We also describe an experimental study of expert programmers debugging real bugs described in real bug reports for a real Java application. We found that PFIS' performance was close to aggregated human decisions as to where to navigate, and was significantly better than individual programmers' decisions.
Joseph Lawrance, Rachel K. E. Bellamy, Margaret M. Burnett, Kyle Rector
CHI3
2008 Testing vs. code inspection vs. what else?: male and female end users' debugging strategies
abstract
Little is known about the strategies end-user programmers use in debugging their programs, and even less is known about gender differences that may exist in these strategies. Without this type of information, designers of end-user programming systems cannot know the "target" at which to aim, if they are to support male and female end-user programmers. We present a study investigating this issue. We asked end-user programmers to debug spreadsheets and to describe their debugging strategies. Using mixed methods, we analyzed their strategies and looked for relationships among participants' strategy choices, gender, and debugging success. Our results indicate that males and females debug in quite different ways, that opportunities for improving support for end-user debugging strategies for both genders are abundant, and that tools currently available to end-user debuggers may be especially deficient in supporting debugging strategies used by females.
Neeraja Subrahmaniyan, Laura Beckwith, Valentina Grigoreanu, Margaret M. Burnett, Susan Wiedenbeck, Vaishnavi Narayanan, Karin Bucht, Russell Drummond, Xiaoli Z. Fern
CHI4
2008 Integrating rich user feedback into intelligent user interfaces
abstract
The potential for machine learning systems to improve via a mutually beneficial exchange of information with users has yet to be explored in much detail. Previously, we found that users were willing to provide a generous amount of rich feedback to machine learning systems, and that the types of some of this rich feedback seem promising for assimilation by machine learning algorithms. Following up on those findings, we ran an experiment to assess the viability of incorporating real-time keyword-based feedback in initial training phases when data is limited. We found that rich feedback improved accuracy but an initial unstable period often caused large fluctuations in classifier behavior. Participants were able to give feedback by relying heavily on system communication in order to respond to changes. The results show that in order to benefit from the user's knowledge, machine learning systems must be able to absorb keyword-based rich feedback in a graceful manner and provide clear explanations of their predictions.
Simone Stumpf, Erin Sullivan, Erin Fitzhenry, Ian Oberst, Weng-Keen Wong, Margaret M. Burnett
IUI6
2008 End-user programming in the wild: A field study of CoScripter scripts
abstract
Although a new class of languages has emerged to enable end users to create their own Web applications, little is known about how end-user programmers actually use such languages in the real world. In this paper, we report a field study on over 1400 scripts collected from the Internet which were created by early adopters of CoScripter, a Web macro programming-by-demonstration language. We contrast these Internet scripts with those written by users inside IBM, and describe script usage and re-usage patterns, features used, and users' clever workarounds for features not present in the language. The results show how users grapple with such programming notions as repetition, generalization, and reuse, sometimes inventing their own devices for these. Finally, we discuss the many scripts we found with social implications, whose purposes were to circumvent intended rules, regulations, and usage norm assumptions of a number of Web sites.
Christopher Bogart, Margaret M. Burnett, Allen Cypher, Christopher Scaffidi
VL/HCC2
2008 Can feature design reduce the gender gap in end-user software development environments?
abstract
Recent research has begun to report that female end-user programmers are often more reluctant than males to employ features that are useful for testing and debugging. These earlier findings suggest that, unless such features can be changed in some appropriate way, there are likely to be important gender differences in end-user programmerspsila benefits from these features. In this paper, we compare end-user programmerspsila feature usage in an environment that supports end-user debugging, against an extension of the same environment with two features designed to help ameliorate the effects of low self-efficacy. Our results show ways in which these features affect female versus male enduser programmerspsila self-efficacy, attitudes, usage of testing and debugging features, and performance.
Valentina Grigoreanu, Jill Cao, Todd Kulesza, Christopher Bogart, Kyle Rector, Margaret M. Burnett, Susan Wiedenbeck
VL/HCC6
2007 Mining Interpretable Human Strategies: A Case Study
abstract
This paper focuses on mining human strategies by observing their actions. Our application domain is an HCI study aimed at discovering general strategies used by software users and understanding how such strategies relate to gender and success. We cast this as a sequential pattern discovery problem, where user strategies are manifested as sequential patterns. Problematically, we found that the patterns discovered by standard algorithms were difficult to interpret and provided limited information about high-level strategies. To help interpret the patterns and extract general strategies, we examined multiple ways of clustering the patterns into meaningful groups, which collectively led to interesting findings about user behavior both in terms of gender differences and problem-solving success. As a real-world application of data mining techniques, our work led to the discovery of new strategic patterns that are linked to user success and had not been revealed in more than nine years of manual empirical work. As a case study, our work highlights important research directions for making data mining more accessible to non-experts.
Xiaoli Z. Fern, Chaitanya Komireddy, Margaret M. Burnett
ICDM3
2007 Toward harnessing user feedback for machine learning
abstract
There has been little research into how end users might be able to communicate advice to machine learning systems. If this resource--the users themselves--could somehow work hand-in-hand with machine learning systems, the accuracy of learning systems could be improved and the users' understanding and trust of the system could improve as well. We conducted a think-aloud study to see how willing users were to provide feedback and to understand what kinds of feedback users could give. Users were shown explanations of machine learning predictions and asked to provide feedback to improve the predictions. We found that users had no difficulty providing generous amounts of feedback. The kinds of feedback ranged from suggestions for reweighting of features to proposals for new features, feature combinations, relational features, and wholesale changes to the learning algorithm. The results show that user feedback has the potential to significantly improve machine learning systems, but that learning algorithms need to be extended in several ways to be able to assimilate this feedback.
Simone Stumpf, Vidya Rajaram, Lida Li, Margaret M. Burnett, Thomas G. Dietterich, Erin Sullivan, Russell Drummond, Jon Herlocker
IUI4
2007 On to the Real World: Gender and Self-Efficacy in Excel
abstract
Although there have been a number of studies of end-user software development tasks, few of them have considered gender issues for real end-user developers in real-world environments for end-user programming. In order to be trusted, the results of such laboratory studies must always be re-evaluated with fewer controls, more closely reflecting real-world conditions. Therefore, the research question in this paper is whether the results of a gender HCI controlled study generalize - to real-world end-user developers, in a real-world spreadsheet environment, using a real-world spreadsheet. Our findings are that the concepts revealed by the original laboratory study appear to be quite robust, being demonstrated in multiple ways in this real-world environment.
Laura Beckwith, Derek Inman, Kyle Rector, Margaret M. Burnett
VL/HCC4
2007 Scents in Programs: Does Information Foraging Theory Apply to Program Maintenance?
abstract
During maintenance, professional developers generate and test many hypotheses about program behavior, but they also spend much of their time navigating among classes and methods. Little is known, however, about how professional developers navigate source code and the extent to which their hypotheses relate to their navigation. A lack of understanding of these issues is a barrier to tools aiming to reduce the large fraction of time developers spend navigating source code. In this paper, we report on a study that makes use of information foraging theory to investigate how professional developers navigate source code during maintenance. Our results showed that information foraging theory was a significant predictor of the developers' maintenance behavior, and suggest how tools used during maintenance can build upon this result, simply by adding word analysis to their reasoning systems.
Joseph Lawrance, Rachel K. E. Bellamy, Margaret M. Burnett
VL/HCC3
2007 Explaining Debugging Strategies to End-User Programmers
abstract
There has been little research into how end-user programming environments can provide explanations that could fill a critical information gap for end-user debuggers - help with debugging strategy. To address this need, we designed and prototyped a video-based approach for explaining debugging strategy, and accompanied it with a text-only approach. We then conducted a qualitative empirical study with end-user debuggers. The results reveal the influences of the explanations on end-user debuggers' decision making, how users reacted to the video versus textual media, and the information gaps the explanations closed. The results also reveal issues of particular importance to explanations of this type.
Neeraja Subrahmaniyan, Cory Kissinger, Kyle Rector, Derek Inman, Jared Kaplan, Laura Beckwith, Margaret M. Burnett
VL/HCC7
2006 Supporting end-user debugging: what do users want to know?
abstract
Although researchers have begun to explicitly support end-user programmers' debugging by providing information to help them find bugs, there is little research addressing the right content to communicate to these users. The specific semantic content of these debugging communications matters because, if the users are not actually seeking the information the system is providing, they are not likely to attend to it. This paper reports a formative empirical study that sheds light on what end users actually want to know in the course of debugging a spreadsheet, given the availability of a set of interactive visual testing and debugging features. Our results provide in sights into end-user debuggers' information gaps, and further suggest opportunities to improve end-user debugging systems' support for the things end-user debuggers actually want to know.
Cory Kissinger, Margaret M. Burnett, Simone Stumpf, Neeraja Subrahmaniyan, Laura Beckwith, Sherry Yang 0002, Mary Beth Rosson
AVI2
2006 Tinkering and gender in end-user programmers' debugging
abstract
Earlier research on gender effects with software features intended to help problem-solvers in end-user debugging environments has shown that females are less likely to use unfamiliar software features. This poses a serious problem because these features may be key to helping them with debugging problems. Contrasting this with research documenting males' inclination for tinkering in unfamiliar environments, the question arises as to whether encouraging tinkering with new features would help females overcome the factors, such as low self-efficacy, that led to the earlier results. In this paper, we present an experiment with males and females in an end-user debugging setting, and investigate how tinkering behavior impacts several measures of their debugging success. Our results show that the factors of tinkering, reflection, and self-efficacy, can combine in multiple ways to impact debugging effectiveness differently for males than for females.
Laura Beckwith, Cory Kissinger, Margaret M. Burnett, Susan Wiedenbeck, Joseph Lawrance, Alan F. Blackwell, Curtis R. Cook
CHI3
2006 Scaling a Dataflow Testing Methodology to the MultiparadigmWorld of Commercial Spreadsheets
abstract
Spreadsheets are widely used but often contain faults. Thus, in prior work we presented a dataflow testing methodology for use with spreadsheets, which studies have shown can be used cost-effectively by end-user programmers. To date, however, the methodology has been investigated across a limited set of spreadsheet language features. Commercial spreadsheet environments are multiparadigm languages, utilizing features not accommodated by our prior approaches. In addition, most spreadsheets contain large numbers of replicated formulas that severely limit the efficiency of dataflow testing approaches. We show how to handle these two issues with a new dataflow adequacy criterion and automated detection of areas of replicated formulas, and report results of a controlled experiment investigating the feasibility of our approach
Marc Fisher II, Gregg Rothermel, Tyler Creelan, Margaret M. Burnett
ISSRE4
2006 Pair Collaboration in End-User Debugging
abstract
The problem of dependability in end-user programming is an emerging area of interest. Pair collaboration in end-user software development may offer a way for end users to debug their programs more effectively. While pair programming studies - primarily of computer science students and professionals - report positive outcomes in terms of overall program quality, little is known about specific activities that pairs engage in that lead to those outcomes, or of how the previous results may pertain to end-user programmers. In this paper we analyze protocols of end-user pairs debugging spreadsheets. The results suggest that end-user pairs can achieve rich reasoning, effective planning, and systematic evaluation. Furthermore, end-user pairs provide specific types of mutual support that facilitate the accomplishment of their goals
Thippaya Chintakovid, Susan Wiedenbeck, Margaret M. Burnett, Valentina Grigoreanu
VL/HCC3
2006 Gender Differences in End-User Debugging, Revisited: What the Miners Found
abstract
We have been working to uncover gender differences in the ways males and females problem solve in end-user programming situations, and have discovered differences in males' versus females' use of several debugging features. Still, because this line of investigation is new, knowing exactly what to look for is difficult and important information could escape our notice. We therefore decided to bring data mining techniques to bear on our data, with two aims: primarily, to expand what is known about how males versus females make use of end-user debugging features, and secondarily, to find out whether data mining could bring new understanding to this research, given that we had already studied the data manually using qualitative and quantitative methods. The results suggested several new hypotheses in how males versus females go about end-user debugging tasks, the factors that play into their choices, and how their choices are associated with success
Valentina Grigoreanu, Laura Beckwith, Xiaoli Z. Fern, Sherry Yang 0002, Chaitanya Komireddy, Vaishnavi Narayanan, Curtis R. Cook, Margaret M. Burnett
VL/HCC8
2006 Sharing reasoning about faults in spreadsheets: An empirical study
abstract
Although researchers have developed several ways to reason about the location of faults in spreadsheets, no single form of reasoning is without limitations. Multiple types of errors can appear in spreadsheets, and various fault localization techniques differ in the kinds of errors that they are effective in locating. In this paper, we report empirical results from an emerging system that attempts to improve fault localization for end-user programmers by sharing the results of the reasoning systems found in WYSIWYT and UCheck. By evaluating the visual feedback from each fault localization system, we shed light on where these different forms of reasoning and combinations of them complement - and contradict - one another, and which heuristics can be used to generate the best advice from a combination of these systems
Joseph Lawrance, Robin Abraham, Margaret M. Burnett, Martin Erwig
VL/HCC3
2006 Integrating automated test generation into the WYSIWYT spreadsheet testing methodology
abstract
Spreadsheet languages, which include commercial spreadsheets and various research systems, have had a substantial impact on end-user computing. Research shows, however, that spreadsheets often contain faults. Thus, in previous work we presented a methodology that helps spreadsheet users test their spreadsheet formulas. Our empirical studies have shown that end users can use this methodology to test spreadsheets more adequately and efficiently; however, the process of generating test cases can still present a significant impediment. To address this problem, we have been investigating how to incorporate automated test case generation into our testing methodology in ways that support incremental testing and provide immediate visual feedback. We have used two techniques for generating test cases, one involving random selection and one involving a goal-oriented approach. We describe these techniques and their integration into our testing environment, and report results of an experiment examining their effectiveness and efficiency.
Marc Fisher II, Gregg Rothermel, Darren Brown, Mingming Cao, Curtis R. Cook, Margaret M. Burnett
ACM Trans. Softw. Eng. Methodol.6
2006 Interactive Fault Localization Techniques in a Spreadsheet Environment
abstract
End-user programmers develop more software than any other group of programmers, using software authoring devices such as multimedia simulation builders, e-mail filtering editors, by-demonstration macro builders, and spreadsheet environments. Despite this, there has been only a little research on finding ways to help these programmers with the dependability of the software they create. We have been working to address this problem in several ways, one of which includes supporting end-user debugging activities through interactive fault localization techniques. This paper investigates fault localization techniques in the spreadsheet domain, the most common type of end-user programming environment. We investigate a technique previously described in the research literature and two new techniques. We present the results of an empirical study to examine the impact of two individual factors on the effectiveness of fault localization techniques. Our results reveal several insights into the contributions such techniques can make to the end-user debugging process and highlight key issues of interest to researchers and practitioners who may design and evaluate future fault localization techniques.
Joseph R. Ruthruff, Margaret M. Burnett, Gregg Rothermel
IEEE Trans. Software Eng.2
2005 Effectiveness of end-user debugging software features: are there gender issues?
abstract
Although gender differences in a technological world are receiving significant research attention, much of the research and practice has aimed at how society and education can impact the successes and retention of female computer science professionals-but the possibility of gender issues within software has received almost no attention. If gender issues exist with some types of software features, it is possible that accommodating them by changing these features can increase effectiveness, but only if we know what these issues are. In this paper, we empirically investigate gender differences for end users in the context of debugging spreadsheets. Our results uncover significant gender differences in self-efficacy and feature acceptance, with females exhibiting lower self-efficacy and lower feature acceptance. The results also show that these differences can significantly reduce females' effectiveness.
Laura Beckwith, Margaret M. Burnett, Susan Wiedenbeck, Curtis R. Cook, Shraddha Sorte, Michelle Hastings
CHI2
2005 An empirical study of fault localization for end-user programmers
abstract
End users develop more software than any other group of programmers, using software authoring devices such as e-mail filtering editors, by-demonstration macro builders, and spreadsheet environments. Despite this, there has been little research on finding ways to help these programmers with the dependability of their software. We have been addressing this problem in several ways, one of which includes supporting end-user debugging activities through fault localization techniques. This paper presents the results of an empirical study conducted in an end-user programming environment to examine the impact of two separate factors in fault localization techniques that affect technique effectiveness. Our results shed new insights into fault localization techniques for end-user programmers and the factors that affect them, with significant implications for the evaluation of those techniques.
Joseph R. Ruthruff, Margaret M. Burnett, Gregg Rothermel
ICSE2
2005 Designing Features for Both Genders in End-User Programming Environments
abstract
Previous research has revealed gender differences that impact females' willingness to adopt software features in end users' programming environments. Since these features have separately been shown to help end users problem solve, it is important to female end users' productivity that we find ways to make these features more acceptable to females. In this paper, we draw from our ongoing work with users to help inform our design of theory-based methods for encouraging effective feature usage by both genders. This design effort is the first to begin addressing the gender differences in the ways that people go about problem solving in end-user programming situations.
Laura Beckwith, Shraddha Sorte, Margaret M. Burnett, Susan Wiedenbeck, Thippaya Chintakovid, Curtis R. Cook
VL/HCC3
2005 How Well Do Professional Developers Test with Code Coverage Visualizations? An Empirical Study
abstract
Despite years of availability of testing tools, professional software developers still seem to need better support to determine the effectiveness of their tests. Without improvements in this area, inadequate testing of software seems likely to remain a major problem. To address this problem, industry and researchers have proposed systems that visualize "testedness" for end-user and professional developers. Empirical studies of such systems for end-user programmers have begun to show success at helping end users write more effective tests. Encouraged by this research, we examined the effect that code coverage visualizations have on the effectiveness of test cases that professional software developers write. This paper presents the results of an empirical study conducted using code coverage visualizations found in a commercially available programming environment. Our results reveal how this kind of code coverage visualization impacts test effectiveness, and provide insights into the strategies developers use to test code.
Joseph Lawrance, Steven Clarke, Margaret M. Burnett, Gregg Rothermel
VL/HCC3
2005 Garbage in, Garbage out? An Empirical Look at Oracle Mistakes by End-User Programmers
abstract
End-user programmers, because they are human, make mistakes. However, past research has not considered how visual end-user debugging devices could be designed to ameliorate the effects of mistakes. This paper empirically examines oracle mistakes - mistakes users make about which values are right and which are wrong - to reveal differences in how different types of oracle mistakes impact the quality of visual feedback about bugs. We then consider the implications of these empirical results for designers of end-user software engineering environments.
Amit Phalgune, Cory Kissinger, Margaret M. Burnett, Curtis R. Cook, Laura Beckwith, Joseph R. Ruthruff
VL/HCC3
2005 The impact of software engineering research on modern programming languages
abstract
Software engineering research and programming language design have enjoyed asymbioticrelationship, with traceable impacts since the 1970s, when these areas were first distinguished from one another. This report documents this relationship by focusing on several major features of current programming languages: data and procedural abstraction, types, concurrency, exceptions, and visual programming mechanisms. The influences are determined by tracing references in publications in both fields, obtaining oral histories from language designers delineating influences on them, and tracking cotemporal research trends and ideas as demonstrated by workshop topics, special issue publications, and invited talks in the two fields. In some cases there is conclusive data supporting influence. In other cases, there are circumstantial arguments (i.e., cotemporal ideas) that indicate influence. Using this approach, this study provides evidence of the impact of software engineering research on modern programming language design and documents the close relationship between these two fields.
Barbara G. Ryder, Mary Lou Soffa, Margaret M. Burnett
ACM Trans. Softw. Eng. Methodol.3
2004 Impact of interruption style on end-user debugging
abstract
Although researchers have begun to explicitly support end-user programmers' debugging by providing information to help them find bugs, there is little research addressing the proper mechanism to alert the user to this information. The choice of alerting mechanism can be important, because as previous research has shown, different interruption styles have different potential advantages and disadvantages. To explore impacts of interruptions in the end-user debugging domain, this paper describes an empirical comparison of two interruption styles that have been used to alert end-user programmers to debugging information. Our results show that negotiated-style interruptions were superior to immediate-style interruptions in several issues of importance to end-user debugging, and further suggest that a reason for this superiority may be that immediate-style interruptions encourage different debugging strategies.
T. J. Robertson, Shrinu Prabhakararao, Margaret M. Burnett, Curtis R. Cook, Joseph R. Ruthruff, Laura Beckwith, Amit Phalgune
CHI3
2004 Gender: An Important Factor in End-User Programming Environments?
abstract
A human-centric issue that has not been considered in the design of end-user programming environments is whether gender differences exist that are important to the design of these environments. Ignoring this issue would miss the opportunity of enhancing the effectiveness of end-user programmers by incorporating appropriate mechanisms to support gender-associated differences in decision making, learning, and problem solving. This paper takes a first step toward building a foundation for investigating this issue by surveying gender difference literature from five domains with an eye toward possible implications for end-user programming. We present a taxonomy of this literature, and derive a number of specific issues for each element of the taxonomy (stated as hypotheses). This foundation provides a starting point for organized investigations into issues that may be important for making breakthroughs in the effectiveness of end-user programmers
Laura Beckwith, Margaret M. Burnett
VL/HCC2
2004 Champagne Prototyping: A Research Technique for Early Evaluation of Complex End-User Programming Systems
abstract
Although a variety of evaluation techniques are available to researchers of visual and end-user programming systems, they are primarily suited to evaluation of research systems. It is important to have evaluation techniques suitable for real-world programming environments, in order to satisfy real-world product managers of the usefulness of proposed new features. To help fill this gap, we present a new evaluation technique, based in part on Cognitive Dimensions and Attention Investment, called "Champagne Prototyping". The technique is an early-evaluation technique that is inexpensive to do, yet features the credibility that comes from being based on the real commercial environment of interest, and from working with real users of the environment.
Alan F. Blackwell, Margaret M. Burnett, Simon L. Peyton Jones
VL/HCC2
2004 Rewarding "Good" Behavior: End-User Debugging and Rewards
abstract
Emerging research has sought to bring effective debugging devices to end-user programmers. This research has largely focused on how well such devices bring genuine "functional" rewards to end users. However, emerging models of programming behavior indicate that another, often ignored, type of reward-perceivable rewards-can play an equally vital role in how well debugging devices serve end users. Using an empirically evaluated fault localization device, this paper investigates the impact such perceivable rewards can have on end-user debugging. Our results indicate that perceivable rewards alone can significantly improve the effectiveness and understanding of end users performing debugging tasks.
Joseph R. Ruthruff, Amit Phalgune, Laura Beckwith, Margaret M. Burnett, Curtis R. Cook
VL/HCC4
2003 Harnessing curiosity to increase correctness in end-user programming
abstract
Despite their ability to help with program correctness, assertions have been notoriously unpopular---even with professional programmers. End-user programmers seem even less likely to appreciate the value of assertions; yet end-user programs suffer from serious correctness problems that assertions could help detect. This leads to the following question: can end users be enticed to enter assertions? To investigate this question, we have devised a curiosity-centered approach to eliciting assertions from end users, built on a surprise-explain-reward strategy. Our follow-up work with end-user participants shows that the approach is effective in encouraging end users to enter assertions that help them find errors.
Margaret M. Burnett, Laura Beckwith, Orion Granatir, Ledah Casburn, Curtis R. Cook, Mike Durham, Gregg Rothermel
CHI2
2003 A user-centred approach to functions in Excel
abstract
We describe extensions to the Excel spreadsheet that integrate userdefined functions into the spreadsheet grid, rather than treating them as a "bolt-on". Our first objective was to bring the benefits of additional programming language features to a system that is often not recognised as a programming language. Second, in a project involving the evolution of a well-established language, compatibility with previous versions is a major issue, and maintaining this compatibility was our second objective. Third and most important, the commercial success of spreadsheets is largely due to the fact that many people find them more usable than programming languages for programming-like tasks. Thus, our third objective (with resulting constraints) was to maintain this usability advantage.
Simon L. Peyton Jones, Alan F. Blackwell, Margaret M. Burnett
ICFP3
2003 End-User Software Engineering with Assertions in the Spreadsheet Paradigm
abstract
There has been little research on end-user program development beyond the activity of programming. Devising ways to address additional activities related to end-user program development may be critical, however, because research shows that a large proportion of the programs written by end users contain faults. Toward this end, we have been working on ways to provide formal "software engineering" methodologies to end-user programmers. This paper describes an approach we have developed for supporting assertions in end-user software, focusing on the spreadsheet paradigm. We also report the results of a controlled experiment, with 59 end-user subjects, to investigate the usefulness of this approach. Our results show that the end users were able to use the assertions to reason about their spreadsheets, and that doing so was tied to both greater correctness and greater efficiency.
Margaret M. Burnett, Curtis R. Cook, Omkar Pendse, Gregg Rothermel, Jay Summet, Christine Wallace
ICSE1
2003 End-User Testing for the Lyee Methodology using the Screen Transition Paradigm and WYSIWYT
Darren Brown, Margaret M. Burnett, Gregg Rothermel
Knowl. Based Syst.2
2003 HCI research regarding end-user requirement specification: a tutorial
Margaret M. Burnett
Knowl. Based Syst.1
2002 Automated test case generation for spreadsheets
abstract
Spreadsheet languages, which include commercial spreadsheets and various research systems, have had a substantial impact on end-user computing. Research shows, however, that spreadsheets often contain faults. Thus, in previous work, we presented a methodology that assists spreadsheet users in testing their spreadsheet formulas. Our empirical studies have shown that this methodology can help end-users test spreadsheets more adequately and efficiently; however, the process of generating test cases can still represent a significant impediment. To address this problem, we have been investigating how to automate test case generation for spreadsheets in ways that support incremental testing and provide immediate visual feedback. We have utilized two techniques for generating test cases, one involving random selection and one involving a goal-oriented approach. We describe these techniques, and report results of an experiment examining their relative costs and benefits.
Marc Fisher II, Mingming Cao, Gregg Rothermel, Curtis R. Cook, Margaret M. Burnett
ICSE5
2002 Test Reuse in the Spreadsheet Paradigm
abstract
Spreadsheet languages are widely used by a variety of end users to perform many important tasks. Despite their perceived simplicity, spreadsheets often contain faults. Furthermore, users modify their spreadsheets frequently, which can render previously correct spreadsheets faulty. To address this problem, we previously introduced a visual approach by which users can systematically test their spreadsheets, see where new tests are required after changes, and request automated generation of potentially useful test inputs. To date, however, this approach has not taken advantage of previously developed test cases, which means that users of the approach cannot benefit, when re-testing following changes, from prior testing efforts. We have therefore been investigating ways to add support for test re-use into our spreadsheet testing methodology. In this paper we present a test re-use strategy for spreadsheets, and the algorithms that implement it, and describe their integration into our spreadsheet testing methodology. We report results of a case study examining the application of this strategy.
Marc Fisher II, Dalai Jin, Gregg Rothermel, Margaret M. Burnett
ISSRE4
2002 Adding Apples and Oranges
Martin Erwig, Margaret M. Burnett
PADL2
2002 Appendices A-D: A scalable method for deductive generalization in the spreadsheet paradigm
abstract
article Appendices A--D: A scalable method for deductive generalization in the spreadsheet paradigm Share on Authors: Margaret Burnett Oregon State University Oregon State UniversityView Profile , Sherry Yang Oregon Institute of Technology Oregon Institute of TechnologyView Profile , Jay Summet Oregon State University Oregon State UniversityView Profile Authors Info & Claims ACM Transactions on Computer-Human InteractionVolume 9Issue 4December 2002 pp 1–5https://doi.org/10.1145/586081.586082Published:01 December 2002 0citation420DownloadsMetricsTotal Citations0Total Downloads420Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Margaret M. Burnett, Sherry Yang 0002, Jay Summet
ACM Trans. Comput. Hum. Interact.1
2002 A scalable method for deductive generalization in the spreadsheet paradigm
abstract
In this paper, we present an efficient method for automatically generalizing programs written in spreadsheet languages. The strategy is to do generalization through incremental analysis of logical relationships among concrete program entities from the perspective of a particular computational goal. The method uses deductive dataflow analysis with algebraic back-substitution rather than inference with heuristics, and there is no need for generalization-related dialog with the user. We present the algorithms and their time complexities and show that, because the algorithms perform their analyses incrementally, on only the on-screen program elements rather than on the entire program, the method is scalable. Performance data is presented to help demonstrate the scalability.
Margaret M. Burnett, Sherry Yang 0002, Jay Summet
ACM Trans. Comput. Hum. Interact.1
2002 Testing Homogeneous Spreadsheet Grids with the "What You See Is What You Test" Methodology
abstract
Although there has been recent research into ways to design environments that enable end users to create their own programs, little attention has been given to helping these end users systematically test their programs. To help address this need in spreadsheet systems (the most widely used type of end-user programming language), we previously introduced a visual approach to systematically testing individual cells in spreadsheet systems. However, the previous approach did not scale well in the presence of largely homogeneous grids, which introduce problems somewhat analogous to the array-testing problems of imperative programs. We present two approaches to spreadsheet testing that explicitly support such grids. We present the algorithms, time complexities, and performance data comparing the two approaches. This is part of our continuing work to bring to end users at least some of the benefits of formalized notions of testing without requiring knowledge of testing beyond a naive level.
Margaret M. Burnett, Andrei Sheretov, Gregg Rothermel
IEEE Trans. Software Eng.1
2001 Incorporating Incremental Validation and Impact Analysis into Spreadsheet Maintenance: An Empirical Study
abstract
Spreadsheets are among the most common form of software in use today. Unlike more traditional forms of software however, spreadsheets are created and maintained by end users with little or no programming experience. As a result, a high percentage of these "programs" contain errors. Unfortunately, software engineering research has for the most part ignored this problem. We have developed a methodology that is designed to aid end users in developing, testing, and maintaining spreadsheets. The methodology communicates testing information and information about the impact of cell changes to users in a manner that does not require an understanding of formal testing theory or the behind the scenes mechanisms. The paper presents the results of an empirical study that shows that, during maintenance, end users using our methodology were more accurate in making changes and did a significantly better job of validating their spreadsheets than end users without the methodology.
Vijay B. Krishna, Curtis R. Cook, Daniel Keller, Joshua Cantrell, Christine Wallace, Margaret M. Burnett, Gregg Rothermel
ICSM6
2001 Forms/3: A first-order visual language to explore the boundaries of the spreadsheet paradigm
abstract
Although detractors of functional programming sometimes claim that functional programming is too difficult or counter-intuitive for most programmers to understand and use, evidence to the contrary can be found by looking at the popularity of spreadsheets. The spreadsheet paradigm, a first-order subset of the functional programming paradigm, has found wide acceptance among both programmers and end users. Still, there are many limitations with most spreadsheet systems. In this paper, we discuss language features that eliminate several of these limitations without deviating from the first-order, declarative evaluation model. The language used to illustrate these features is a research language called Forms/3. Using Forms/3, we show that procedural abstraction, data abstraction and graphics output can be supported in the spreadsheet paradigm. We show that, with the addition of a simple model of time, animated output and GUI I/O also become viable. To demonstrate generality, we also present an animated Turing machine simulator programmed using these features. Throughout the paper, we combine our discussion of the programming language characteristics with how the language features prototyped in Forms/3 relate to what is known about human effectiveness in programming.
Margaret M. Burnett, John Atwood, Rebecca Walpole Djang, James Reichwein, Herkimer J. Gottfried, Sherry Yang 0002
J. Funct. Program.1
2001 A methodology for testing spreadsheets
abstract
Spreadsheet languages, which include commercial spreadsheets and various research systems, have had a substantial impact on end-user computing. Research shows, however, that spreadsheets often contain faults; thus, we would like to provide at least some of the benefits of formal testing methodologies to the creators of spreadsheets. This article presents a testing methodology that adapts data flow adequacy criteria and coverage monitoring to the task of testing spreadsheets. To accommodate the evaluation model used with spreadsheets, and the interactive process by which they are created, our methodology is incremental. To accommodate the users of spreadsheet languages, we provide an interface to our methodology that does not require an understanding of testing theory. We have implemented our testing methodology in the context of the Forms/3 visual spreadsheet language. We report on the methodology, its time and space costs, and the mapping from the testing strategy to the user interface. In an empirical study, we found that test suites created according to our methodology detected, on average, 81% of the faults in a set of faulty spreadsheets, significantly outperforming randomly generated test suites.
Gregg Rothermel, Margaret M. Burnett, Christopher DuPuis, Andrei Sheretov
ACM Trans. Softw. Eng. Methodol.2
2000 WYSIWYT testing in the spreadsheet paradigm: an empirical evaluation
abstract
Is it possible to achieve some of the benefits of formal testing within the informal programming conventions of the spreadsheet paradigm? We have been working on an approach that attempts to do so via the development of a testing methodology for this paradigm. Our “What You See Is What You Test” (WYSIWYT) methodology supplements the convention by which spreadsheets provide automatic immediate visual feedback about values by providing automatic immediate visual feedback about “testedness”. In previous work we described this methodology; in this paper, we present empirical data about the methodology's effectiveness. Our results show that the use of the methodology was associated with significant improvement in testing effectiveness and efficiency even with no training on the theory of testing or test adequacy that the model implements. These results may be due at least in part to the fact that use of the methodology was associated with a significant reduction in overconfidence.
Karen J. Rothermel, Curtis R. Cook, Margaret M. Burnett, Justin Schonfeld, Thomas R. G. Green, Gregg Rothermel
ICSE3
2000 Exception Handling in the Spreadsheet Paradigm
abstract
Exception handling is widely regarded as a necessity in programming languages today and almost every programming language currently used for professional software development supports some form of it. However, spreadsheet systems, which may be the most widely used type of "programming language" today in terms of number of users using it to create "programs" (spreadsheets), have traditionally had only extremely limited support for exception handling. Spreadsheet system users range from end users to professional programmers and this wide range suggests that an approach to exception handling for spreadsheet systems needs to be compatible with the equational reasoning model of spreadsheet formulas, yet feature expressive power comparable to that found in other programming languages. We present an approach to exception handling for spreadsheet system users that is aimed at this goal. Some of the features of the approach are new; others are not new, but their effects on the programming language properties of spreadsheet systems have not been discussed before in the literature. We explore these properties, offer our solutions to problems that arise with these properties, and compare the functionality of the approach with that of exception handling approaches in other languages.
Margaret M. Burnett, Anurag Agrawal, Pieter van Zee
IEEE Trans. Software Eng.1
1998 Challenges and Oppurtunities Visual Programming Languages Bring to Programming Language Research
Margaret M. Burnett
CC1
1998 What You See Is What You Test: A Methodology for Testing Form-Based Visual Programs
abstract
Form-based visual programming languages, which include commercial spreadsheets and various research systems, have had a substantial impact on end-user computing. Research shows, however, that form-based visual programs often contain faults. We would like to provide at least some of the benefits of formal testing methodologies to the creators of these programs. This paper presents a testing methodology for form-based visual programs. To accommodate the evaluation model used with these programs, and the interactive process by which they are created, our methodology is validation driven and incremental. To accommodate the users of these languages, We provide an interface to the methodology that does not require an understanding of testing theory. We discuss our implementation of this methodology and empirical results achieved in its use.
Gregg Rothermel, Christopher DuPuis, Margaret M. Burnett
ICSE4
1998 Graphical Definitions: Expanding Spreadsheet Languages Through Direct Manipulation and Gestures
abstract
In the past, attempts to extend the spreadsheet paradigm to support graphical objects, such as colored circles or user-defined graphical types, have led to approaches featuring either a direct way of creating objects graphically or strong compatibility with the spreadsheet paradigm, but not both. This inability to conveniently go beyond numbers and strings without straying outside the spreadsheet paradigm has been a limiting factor in the applicability of spreadsheet languages. In this article we present graphical definitions, an approach that removes this limitation, allowing both simple and complex graphical objects to be programmed directly using direct manipulation and gestures, in a manner that fits seamlessly within the spreadsheet paradigm. We also describe an empirical study, in which subjects programmed such objects faster and with fewer errors using this approach than when using a traditional approach to formula specification. Because the approach is expressive enough to be used with both built-in and user-defined types, it allows the directness of demonstrational and spreadsheet techniques to be used in programming a wider range of applications than has been possible before.
Margaret M. Burnett, Herkimer J. Gottfried
ACM Trans. Comput. Hum. Interact.1
1997 Does Continuous Visual Feedback Aid Debugging in Direct-Manipulation Programming Systems?
abstract
Continuous visual feedback is becoming a common feature in direct-manipulation programming systems of all klndsfrom demonstrational macro builders to spreadsheet packages to visual programming languages featuring direct manipulation.But does continuous visual feedback actually help in the domain of programming?There has been little investigation of this question, and what evidence there is from related domains points in conflicting directions.To advance what is known about this issue, we conducted an empirical study to determine whether the inclusion of continuous visual feedback into a direct-manipulation programming system helps with one particular task: debugging.Our results were that although continuous visual feedback did not significantly help with debugging in general, it did significantly help with debugging in some circumstances.Our results also indicate three factors that may help determine those circumstances.
E. M. Wilcox, John Atwood, Margaret M. Burnett, Jonathan J. Cadiz, Curtis R. Cook
CHI3
1997 Testing strategies for form-based visual programs
abstract
Form based visual programming languages, which include electronic spreadsheets and a variety of research systems, have had a substantial impact on end user computing. Research shows that form based visual programs often contain faults, and that their creators often have unwarranted confidence in the reliability of their programs. Despite this evidence, we find no discussion in the research literature of techniques for testing or assessing the reliability of form based visual programs. The paper addresses this lack. We describe differences between the form based and imperative programming paradigms, and discuss effects these differences have on strategies for testing form based programs. We then present several test adequacy criteria for form based programs, and illustrate their application. We show that an analogue to the traditional "all-uses" dataflow test adequacy criterion is well suited for code based testing of form based visual programs: it provides important error detection ability, and can be applied more easily to form based programs than to imperative programs.
Gregg Rothermel, Margaret M. Burnett
ISSRE3