Christopher Bogart

dblp:83/4898 · also Chris Bogart, Christopher A. Bogart · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0001-8581-115XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 33 · 7 first-author · 20 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Teaching Debugging in Cloud-Native Application Development: A Course Intervention Study
abstract
We report on a debugging intervention introduced in the Cloud Native course---part of the AI Technicians program, a collaboration between the U.S. Army's AI Integration Center (AI2C) and Carnegie Mellon University. Learners take an introductory programming course prior to Cloud Native, yet their debugging skills did not transfer to this more advanced context. Across two cohorts, students consistently struggled to identify root causes of issues they faced. We describe the intervention we designed to address this gap and the instruments we are using to evaluate its impact.
Michael Sambou, Can Kultur, Christopher Bogart, Jaromír Savelka
ITiCSE (2)3
2026 Partnering with Community College Faculty to Co-Design Intelligent Tutoring Systems for Cybersecurity Workforce Training
abstract
This experience report describes a partnership between community college faculty and learning scientists to co-design Intelligent Tutoring Systems (ITSs) addressing challenges in cybersecurity workforce training. Our co-design approach combined collaborative reflection on student difficulties from prior course offerings with systematic curricular analysis to identify high-impact intervention points. We targeted two challenge areas: strengthening students' ability to contrast key cybersecurity taxonomies, and providing realistic hands-on training without costly infrastructure. The resulting ITSs include: one employing exercises that scaffold comparison of conceptual categories, and another using lightweight simulations to provide experiential learning while circumventing typical cost and time overhead. Both systems incorporate instructional principles grounded in learning science research, including evidence-based features associated with ITS efficacy such as timely hints and feedback. Through iterative classroom deployment and refinement—including adding task-loop adaptivity to offer repeated practice until mastery—we observed encouraging learning outcomes, alongside insights into mitigating ''gaming the system'' behaviors. We detail our co-design process and formative evaluations—procedures, outcomes, and cautious interpretation due to the limited number of consented learners—and share lessons learned to inform scalable ITS development for cybersecurity workforce training in resource-constrained settings.
Marshall An, Mahboobeh Mehrvarz, Leah Teffera, Matthew Kisow, Bruce M. McLaren, Christopher Bogart
SIGCSE (1)6
2025 Auto-Grader Feedback Utilization and Its Impacts: An Observational Study Across Five Community Colleges
abstract
Automated grading systems, or auto-graders, have become ubiquitous in programming education, and the way they generate feedback has become increasingly automated as well. However, there is insufficient evidence regarding auto-grader feedback's effectiveness in improving student learning outcomes, in a way that differentiates students who utilized the feedback and students who did not. In this study, we fill this critical gap. Specifically, we analyze students' interactions with auto-graders in an introductory Python programming course, offered at five community colleges in the United States. Our results show that students checking the feedback more frequently tend to get higher scores from their programming assignments overall. Our results also show that a submission that follows a student checking the feedback tends to receive a higher score than a submission that follows a student ignoring the feedback. Our results provide evidence on auto-grader feedback's effectiveness, encourage their increased utilization, and call for future work to continue their evaluation in this age of automation
Adam Zhang, Heather Burte, Jaromír Savelka, Christopher Bogart, Majd F. Sakr
CSEDU (1)4
2025 Are Students' Evaluations of Auto-Graders Biased by Their Grades?
Jaromír Savelka, Heather Burte, Christopher Bogart, Seth Copen Goldstein, Majd F. Sakr
EC-TEL (2)4
2025 AI Technicians: Developing Rapid Occupational Training Methods for a Competitive AI Workforce
abstract
The accelerating pace of developments in Artificial Intelligence (AI) and the increasing role that technology plays in society necessitates substantial changes in the structure of the workforce. Besides scientists and engineers, there is a need for a very large workforce of competent AI technicians (i.e., maintainers, integrators) and users (i.e., operators). As traditional 4-year and 2-year degree-based education cannot fill this quickly opening gap, alternative training methods have to be developed. We present the results of the first four years of the AI Technicians program which is a unique collaboration between the U.S. Army's Artificial Intelligence Integration Center (AI2C) and Carnegie Mellon University to design, implement and evaluate novel rapid occupational training methods to create a competitive AI workforce at the technicians level. Through this multi-year effort we have already trained 59 AI Technicians. A key observation is that ongoing frequent updates to the training are necessary as the adoption of AI in the U.S. Army and within the society at large is evolving rapidly. A tight collaboration among the stakeholders from the army and the university is essential for successful development and maintenance of the training for the evolving role. Our findings can be leveraged by large organizations that face the challenge of developing a competent AI workforce as well as educators and researchers engaged in solving the challenge.
Jaromír Savelka, Can Kultur, Arav Agarwal, Christopher Bogart, Heather Burte, Adam Zhang, Majd F. Sakr
SIGCSE (1)4
2025 Measuring SES-related traits relating to technology usage: Two validated surveys
abstract
Abstract Software producers are now recognizing the importance of improving their products’ suitability for diverse populations, but little attention has been given to measurements to shed light on products’ suitability to individuals below the median s ocio e conomic s tatus (SES)—who, by definition, make up half the population. To enable software practitioners to attend to both lower- and higher-SES individuals, this paper provides two new surveys that together can facilitate measuring how well a software product serves socioeconomically diverse populations. The first survey (SES-Subjective) is who-oriented: it measures who their potential or current users are in terms of their subjective SES (perceptions of their SES). The second survey (SES-Facets) is why-oriented: it collects individuals’ values for an evidence-based set of facet values (individual traits) that (1) statistically differ by SES and (2) affect how an individual works and problem-solves with software products. The surveys’ design goal is worldwide applicability, but as a first step, here we empirically validated both these surveys with deployments at University A and University B (464 and 522 responses, respectively), which showed reliability of both the surveys in a US context. Our results also statistically agree with both ground truth data on respondents’ socioeconomic statuses and with predictions from foundational literature. Finally, we explain how the pair of surveys can be uniquely actionable by software practitioners, such as in requirements gathering, debugging, quality assurance activities, maintenance activities, and fulfilling legal reporting requirements such as those being drafted by various governments for AI-powered software.
Chimdi Chikezie, Pannapat Chanpaisaeng, Puja Agarwal, Bhavika Madhwani, Rudrajit Choudhuri, Andrew Anderson 0002, Prisha Velhal, Patricia Morreale, Christopher Bogart, Anita Sarma, Margaret M. Burnett
Empir. Softw. Eng.10
2025 Intersectional HCI on a Budget: An Analytical Approach Powered by Types
abstract
Intersectional HCI recognizes that humans' interconnected social identities shape their experiences with technology. However, intersectional HCI requires extensive resources, such as access to intersectional populations, which many HCI practitioners may lack. For these practitioners, we present an analytical approach to bring intersectional lenses to HCI practices. The approach uses types—not at the level of identities, but at the level of personal traits drawn from foundational research. We first formally prove that certain analytical methods for detecting inclusivity issues can be meaningfully composed to provide equitable consideration of typically overlooked populations; then present four design use-cases to illustrate what the approach brings to HCI practices; and then empirically investigated one of the four use-cases with 24 HCI participants. Results show that practitioners using the compositional approach detected even more intersectional inclusivity problems than those using a complementary intersectional approach.
Abrar Fallatah, Md Montaser Hamid, Fatima A. Moussaoui, Chimdi Chikezie, Martin Erwig, Christopher Bogart, Anita Sarma, Margaret M. Burnett
Int. J. Hum. Comput. Interact.6
2024 Generating Situated Reflection Triggers About Alternative Solution Paths: A Case Study of Generative AI for Computer-Supported Collaborative Learning
Atharva Naik, Jessica Ruhan Yin, Anusha Kamath, Qianou Ma, Sherry Tongshuang Wu, R. Charles Murray, Christopher Bogart, Majd F. Sakr, Carolyn P. Rosé
AIED (1)7
2024 Leveraging Intelligent Tutoring Systems to Enhance Project-Based Learning in Workforce Training at Community Colleges
Marshall An, Leah Teffera, Mahboobeh Mehrvarz, Bruce Li, Christopher Bogart, Majd F. Sakr, Bruce M. McLaren
EC-TEL (2)5
2024 Examining the Trade-Offs Between Simplified and Realistic Coding Environments in an Introductory Python Programming Class
Huy Anh Nguyen, Christopher Bogart, Jaromír Savelka, Adam Zhang, Majd F. Sakr
EC-TEL (1)2
2024 Course Delivery Methods, Student Success, and Self-efficacy in Introductory Programming
abstract
Self-efficacy has been claimed to be a predictor of students' motivation and learning [1]. It has been found to be sensitive to students' success, and to affect their academic achievement. In the CS/IT education context, where the drop rates are high, it is important that students not only gain knowledge and skills, but also self-efficacy, so that they persist in the program. In this study, we investigate 602 students taking an introductory Python course via different delivery methods: (i) traditional in-person; (ii) cohort in-person; (iii) synchronous online; and (iv) asynchronous online. Although modality predicted retention and success, we found no apparent links among learning, student retention, and self-efficacy. However we found evidence that cohort learning may in particular help struggling students catch up with their peers.
Christopher Bogart, Can Kultur, Eric Keylor, Jaromír Savelka, Majd F. Sakr
ITiCSE (2)1
2024 Designing Modular Auto-graded Programming Projects
abstract
In this poster we propose an approach to designing auto-graded programming course projects that are modular and easily manageable by an instructor. Based on our experiences with the Sail() platform which supports auto-grading and feedback generation in multiple contexts, we design the approach to overcome the challenges we observed. The approach is especially focused on designing projects that can be utilized by multiple instructors who may have various scopes or students with varying backgrounds. The approach enables differentiated learning-thereby improving learning experiences and outcomes. We also discuss challenges of using such a modular approach to auto-graded projects.
Can Kultur, Jaromír Savelka, Christopher Bogart, Majd F. Sakr
ITiCSE (2)3
2024 What Factors Influence Persistence in Project-based Programming Courses at Community Colleges?
abstract
The rapid adoption of emergent technologies is creating significant shortfall in the CS/IT workforce. With not enough students in the educational pipeline to meet the forthcoming demand over the next decade, community colleges are making the effort to train confident, knowledgeable, and self-driven workers in this field. Project-based learning (PBL) has been shown to be effective for these ends, but it poses distinct challenges in resource-limited community college contexts since it may require more time, preparation, and motivation than other teaching modalities, from both the student and the instructor. We studied fifteen sections of an introductory project-based Python course taught at six community colleges, investigating several features of PBL theorized to be particular barriers to student persistence, particularly among women and other identities traditionally underrepresented in technical fields. We describe successes and challenges faced by students in these areas and suggest implications for project-based learning curriculum and platform design.
Christopher Bogart, Marshall An, Eric Keylor, Pawanjeet Singh, Jaromír Savelka, Majd F. Sakr
SIGCSE (1)1
2024 Programming Plagiarism Detection with Learner Data
abstract
Courses with programming assignments have long faced the issue of academic integrity violations (AIV) where cheating could harm the outcome of student learning. Checking code similarity in students' final submissions is a common way to mitigate this issue. But this single analysis is insufficient as 1) students can refactor their code to evade the check, 2) mere code similarity may not be strong enough evidence to support an AIV case, particularly for simpler assignments that may have similar solutions, and 3) code similarity cannot reveal much about the actual circumstances and behaviors of plagiarism. Due to the lack of supporting data or tools, many educators either abandon solving these challenges or rely on manual approaches that are not feasible at scale. In this paper, we propose a workflow to solve the above challenges for large programming classes by providing supporting evidence of cheating with additional learner data: detailed submission timelines with scores and source code. Running this workflow in a large advanced programming course over several years has helped us identify many cheating cases effectively and efficiently.
Yifan Song 0007, Yuanxin Wang 0001, Marshall An, Christopher Bogart, Majd F. Sakr
SIGCSE (2)4
2024 Assessing the Efficacy of Goal-Based Scenarios in Scaling AI Literacy for Non-Technical Learners
abstract
AI's pervasive role in various fields highlights the imperative for the workforce to adeptly leverage its potential. While numerous courses cater to developers, there exists a discernible void for the wider community of AI users. To address this, our study introduces 'AI User'-a suite of interactive modules hosted on the Sail() platform, designed specifically for non-technical individuals utilizing Goal-Based Scenario (GBS) learning. We conducted a controlled experiment to ascertain whether GBS offers superior learning gains in AI literacy compared to traditional deliberate practice using multiple choice questions.
Ying-Jui Tseng, Ruiwei Xiao, Christopher Bogart, Jaromír Savelka, Majd F. Sakr
SIGCSE (2)3
2023 Large Language Models (GPT) Struggle to Answer Multiple-Choice Questions About Code
Jaromír Savelka, Arav Agarwal, Christopher Bogart, Majd F. Sakr
CSEDU (2)3
2023 Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses
abstract
This paper studies recent developments in large language models’ (LLM) abilities to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. The emergence of ChatGPT resulted in heated debates of its potential uses (e.g., exercise generation, code explanation) as well as misuses in programming classes (e.g., cheating). Recent studies show that while the technology performs surprisingly well on diverse sets of assessment instruments employed in typical programming classes the performance is usually not sufficient to pass the courses. The release of GPT-4 largely emphasized notable improvements in the capabilities related to handling assessments originally designed for human test-takers. This study is the necessary analysis in the context of this ongoing transition towards mature generative AI systems. Specifically, we report the performance of GPT-4, comparing it to the previous generations of GPT models, on three Python courses with assessments ranging from simple multiple-choice questions (no code involved) to complex programming projects with code bases distributed into multiple files (599 exercises overall). Additionally, we analyze the assessments that were not handled well by GPT-4 to understand the current limitations of the model, as well as its capabilities to leverage feedback provided by an auto-grader. We found that the GPT models evolved from completely failing the typical programming class’ assessments (the original GPT-3) to confidently passing the courses with no human involvement (GPT-4). While we identified certain limitations in GPT-4’s handling of MCQs and coding exercises, the rate of improvement across the recent generations of GPT models strongly suggests their potential to handle almost any type of assessment widely used in higher education programming courses. These findings could be leveraged by educators and institutions to adapt the design of programming assessments as well as to fuel the necessary discussions into how programming classes should be updated to reflect the recent technological developments. This study provides evidence that programming instructors need to prepare for a world in which there is an easy-to-use widely accessible technology that can be utilized by learners to collect passing scores, with no effort whatsoever, on what today counts as viable programming knowledge and skills assessments.
Jaromír Savelka, Arav Agarwal, Marshall An, Christopher Bogart, Majd F. Sakr
ICER (1)4
2023 Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses?
abstract
We evaluated the capability of generative pre-trained transformers (GPT), to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. Discussions of potential uses (e.g., exercise generation, code explanation) and misuses (e.g., cheating) of this emerging technology in programming education have intensified, but to date there has not been a rigorous analysis of the models' capabilities in the realistic context of a full-fledged programming course with diverse set of assessment instruments. We evaluated GPT on three Python courses that employ assessments ranging from simple multiple-choice questions (no code involved) to complex programming projects with code bases distributed into multiple files (599 exercises overall). Further, we studied if and how successfully GPT models leverage feedback provided by an auto-grader. We found that the current models are not capable of passing the full spectrum of assessments typically involved in a Python programming course (<70% on even entry-level modules). Yet, it is clear that a straightforward application of these easily accessible models could enable a learner to obtain a non-trivial portion of the overall available score (>55%) in introductory and intermediate courses alike. While the models exhibit remarkable capabilities, including correcting solutions based on auto-grader's feedback, some limitations exist (e.g., poor handling of exercises requiring complex chains of reasoning steps). These findings can be leveraged by instructors wishing to adapt their assessments so that GPT becomes a valuable assistant for a learner as opposed to an end-to-end solution.
Jaromír Savelka, Arav Agarwal, Christopher Bogart, Yifan Song 0007, Majd F. Sakr
ITiCSE (1)3
2022 Cheating Detection in Online Assessments via Timeline Analysis
abstract
The potential for academic integrity violations increases in online courses and instructors must place extra attention on academic integrity, since cheating techniques and costs are different than in the physical classroom. Although students are less supervised and able to study in a self-paced mode in online learning, unauthorized collaboration is still considered to be a serious integrity violation. However, online learning platforms have the advantage that they may capture detailed timelines of student activity. Analysis of these can enable instructors to detect many patterns of collaboration, e.g., working on assessments together, or copying solutions from unauthorized web pages. In this paper, we describe detection methods for several common patterns of alignment between work timelines of pairs of students, and these patterns' relationship with corroborative evidence such as similar answers and unusually fast completion times. We describe data collection necessary to apply the timeline analysis technique to weekly quiz assessments and project submissions, and discuss the strength of evidence the technique can provide in different situations. We have been applying these techniques in an online project-based course over several years, and it has helped instructors to successfully identify potential cheating cases.
Jiameng Du, Yifan Song 0007, Mingxiao An, Marshall An, Christopher Bogart, Majd F. Sakr
SIGCSE (1)5
2021 A Thematic Summarization Dashboard for Navigating Student Reflections at Scale
Yuya Asano, Sreecharan Sankaranarayanan, Majd F. Sakr, Christopher Bogart
ICCE4
2021 Are Working Habits Different Between Well-Performing and at-Risk Students in Online Project-Based Courses?
abstract
We analyze differences in working habits between well-performing and at-risk students using highly-granular data collected from two semesters of an online project-based, upper-level course on cloud computing at a US institution of higher education. Such differentiating metrics may provide deeper insights than interim grades, which are oftentimes the only quantifiable data that is captured and available to an instructor as a proxy for students' learning. Interim grades provide little insight into students' broader work habits and may mask unsustainable learning strategies that result in shallow learning or quickly-forgotten skills/knowledge. The adoption of technology-enhanced learning tools for course delivery, automatic feedback, and grading enable data-informed insight and reflection into students' working habits. This data could allow the detection of early signs of under-prepared students or students in crisis. We empirically assess what working habits, if any, differ among well-performing and at-risk students. From clickstream and other activity data, we derive 22 metrics such as time spent reading project write-ups, timing of starting and finishing work, or break-taking. We also calculate two measures of consistency of each metric measured by a coefficient of variance and a variance of ranking over the semester as well as outlier behavior of a student. Using Z-test and Kolmogorov-Smirnov test, we confirm differences in multiple behavior patterns. Notably, our data suggest that well-performing students start and finish working on a project earlier than at-risk students but they also tend to have fewer submissions which indicate they are more thoughtful about feedback.
Mingxiao An, Jaromír Savelka, Christopher Bogart, Majd F. Sakr
ITiCSE (1)5
2021 Combining Collaborative Reflection based on Worked-Out Examples with Problem-Solving Practice: Designing Collaborative Programming Projects for Learning at Scale
abstract
Computer science pedagogy has overwhelmingly favored problem-solving practice over methods of engagement like worked-out example study especially in advanced classes. This is due to the belief that while these alternative methods may improve student conceptual learning, they may leave them less able to perform on authentic problem-solving tasks from a lack of hands-on practice. In this paper, we perform a direct comparison of this trade-off in a synchronous collaborative programming project by adjusting the boundary between problem-solving and collaborative reflection based on a worked-out example while keeping the total time on task constant. We find that the more time students spent on worked example study, the more was the observed improvement in the pre- to post-test scores with no significant difference in performance on a subsequent problem-solving task. These results, therefore, challenge the dominant place of problem-solving practice in the advanced curricular context and inform the design of collaborative programming projects at scale.
Sreecharan Sankaranarayanan, Siddharth Reddy Kandimalla, Christopher Bogart, R. Charles Murray, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
L@S3
2021 Tracing Vulnerable Code Lineage
abstract
This paper presents results from the MSR 2021 Hackathon. Our team investigates files/projects that contain known security vulnerabilities and how widespread they are throughout repositories in open source software. These security vulnerabilities can potentially be propagated through code reuse even when the vulnerability is fixed in different versions of the code. We utilize the World of Code [1] infrastructure to discover file-level duplication of code from a nearly complete collection of open source software. This paper describes a method and set of tools to find all open source projects that use known vulnerable files and any previous revisions of those files.
David Reid, Kalvin Eng, Christopher Bogart, Adam Tutko
MSR3
2021 World of code: enabling a research workflow for mining and analyzing the universe of open source VCS data
Yuxing Ma, Tapajit Dey, Christopher Bogart, Sadika Amreen, Marat Valiev, Adam Tutko, David Kennard, Russell Zaretzki, Audris Mockus
Empir. Softw. Eng.3
2021 "They Can Only Ever Guide": How an Open Source Software Community Uses Roadmaps to Coordinate Effort
abstract
Unlike in commercial software development, open source software (OSS) projects do not generally have managers with direct control over how developers spend their time, yet for projects with large, diverse sets of contributors, the need exists to focus and steer development in a particular direction in a coordinated way. This is especially important for "infrastructure" projects, such as critical libraries and programming languages that many other people depend on. Some projects have taken the approach of borrowing planning tools that originated in commercial development, despite the fact that these techniques were designed for very different contexts, e.g. strong top-down control and profit motives. Little research has been done to understand how these practices are adapted to a new context. In this paper, we examine the Rust project's use of roadmaps: how has an important OSS infrastructure project adapted an inherently top-down tool to the freewheeling world of OSS? We find that because Rust's roadmaps are built in part by summarizing what motivated developers most prefer to work on, they are in some ways more a description of the motivated labor available than they are a directive that the community move in a particular direction. They allow the community to avoid wasting time on unpopular proposals by revealing that there will be little help in building them, and encouraging work on popular features by making visible the amount of consensus in those features. Roadmaps generate a collective focus without limiting the full scope of what developers work on: roadmap issues consume proportionally more effort than other issues, but constitute a minority of the work done (i.e issues and pull requests made) by both central and peripheral participants. They also create transparency among and beyond the community into what central contributors' plans are, and allow more rational decision-making by providing a way for evidence about community needs to be linked to decision-making.
Daniel Klug, Christopher Bogart, James D. Herbsleb
Proc. ACM Hum. Comput. Interact.2
2021 When and How to Make Breaking Changes: Policies and Practices in 18 Open Source Software Ecosystems
abstract
Open source software projects often rely on package management systems that help projects discover, incorporate, and maintain dependencies on other packages, maintained by other people. Such systems save a great deal of effort over ad hoc ways of advertising, packaging, and transmitting useful libraries, but coordination among project teams is still needed when one package makes a breaking change affecting other packages. Ecosystems differ in their approaches to breaking changes, and there is no general theory to explain the relationships between features, behavioral norms, ecosystem outcomes, and motivating values. We address this through two empirical studies. In an interview case study, we contrast Eclipse, NPM, and CRAN, demonstrating that these different norms for coordination of breaking changes shift the costs of using and maintaining the software among stakeholders, appropriate to each ecosystem’s mission. In a second study, we combine a survey, repository mining, and document analysis to broaden and systematize these observations across 18 ecosystems. We find that all ecosystems share values such as stability and compatibility, but differ in other values. Ecosystems’ practices often support their espoused values, but in surprisingly diverse ways. The data provides counterevidence against easy generalizations about why ecosystem communities do what they do.
Christopher Bogart, Christian Kästner, James D. Herbsleb, Ferdian Thung
ACM Trans. Softw. Eng. Methodol.1
2020 Agent-in-the-Loop: Conversational Agent Support in Service of Reflection for Learning During Collaborative Programming
Sreecharan Sankaranarayanan, Siddharth Reddy Kandimalla, Sahil Hasan, Haokang An, Christopher Bogart, R. Charles Murray, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
AIED (2)5
2020 Using Productive Collaboration Bursts to Analyze Open Source Collaboration Effectiveness
abstract
Developers of open-source software projects tend to collaborate in bursts of activity over a few days at a time, rather than at an even pace. A project might find its productivity suffering if bursts of activity occur when a key person with the right role or right expertise is not available to participate. Open-source projects could benefit from monitoring the way they orchestrate attention among key developers, finding ways to make themselves available to one another when needed. In commercial software development, Sociotechnical Congruence (STC) has been used as a measure to assess whether coordination among developers is sufficient for a given task. However, STC has not previously been successfully applied to open-source projects, in which some industrial assumptions do not apply: management-chosen targets, mandated steady work hours, and top-down task allocation of inputs and targets. In this work we propose an operationalization of STC for open-source software development. We use temporal bursts of activity as a unit of analysis more suited to the natural rhythms of open-source work, as well as open source analogues of other component measures needed for calculating STC. As an illustration, we demonstrate that open-source development on PyPI projects in GitHub is indeed bursty, that activities in the bursts have topical coherence, and we apply our operationalization of STC. We argue that a measure of socio-technical congruence adapted to open source could provide projects with a better way of tracking how effectively they are collaborating when they come together to collaborate.
Samridhi Choudhary, Christopher Bogart, Carolyn P. Rosé, James D. Herbsleb
SANER2
2020 ALFAA: Active Learning Fingerprint based Anti-Aliasing for correcting developer identity errors in version control systems
Sadika Amreen, Audris Mockus, Russell Zaretzki, Christopher Bogart, Yuxia Zhang
Empir. Softw. Eng.4
2019 A Qualitative Study on Framework Debugging
abstract
Features of frameworks, such as inversion of control and the structure of framework applications, require developers to adjust their programming and debugging strategies as compared to sequential programs. However, the benefits and challenges of framework debugging are not fully understood, and gaining this knowledge could provide guidance in debugging strategies and framework tool design. To gain insight into the framework application debugging process, we performed two human studies investigating how developers fix applications that use a framework API incorrectly. These studies focused on the Android Fragment class and the ROS framework. We analyzed the results of the studies using a mixed-methods approach, using techniques from qualitative approaches. Our analysis found that participants benefited from the structure of frameworks and the pre-made solutions to common problems in the domain. Participants encountered challenges with understanding frame-work abstractions, and had particular difficulty with inversion of control and object protocol issues. When compared to prior work on debugging, these results show that framework applications have unique debugging challenges.
Zack Coker, David Gray Widder, Claire Le Goues, Christopher Bogart, Joshua Sunshine
ICSME4
2019 World of code: an infrastructure for mining the universe of open source VCS data
abstract
Open source software (OSS) is essential for modern society and, while substantial research has been done on individual (typically central) projects, only a limited understanding of the periphery of the entire OSS ecosystem exists. For example, how are tens of millions of projects in the periphery interconnected through technical dependencies, code sharing, or knowledge flows? To answer such questions we a) create a very large and frequently updated collection of version control data for FLOSS projects named World of Code (WoC) and b) provide basic tools for conducting research that depends on measuring interdependencies among all FLOSS projects. Our current WoC implementation is capable of being updated on a monthly basis and contains over 12B git objects. To evaluate its research potential and to create vignettes for its usage, we employ WoC in conducting several research tasks. In particular, we find that it is capable of supporting trend evaluation, ecosystem measurement, and the determination of package usage. We expect WoC to spur investigation into global properties of OSS development leading to increased resiliency of the entire OSS ecosystem. Our infrastructure facilitates the discovery of key technical dependencies, code flow, and social networks that provide the basis to determine the structure and evolution of the relationships that drive FLOSS activities and innovation.
Yuxing Ma, Christopher Bogart, Sadika Amreen, Russell Zaretzki, Audris Mockus
MSR2
2018 When Optimal Team Formation Is a Choice - Self-selection Versus Intelligent Team Formation Strategies in a Large Online Project-Based Course
Sreecharan Sankaranarayanan, Cameron Dashti, Christopher Bogart, Xu Wang 0016, Majd F. Sakr, Carolyn P. Rosé
AIED (1)3
2016 How to break an API: cost negotiation and community values in three software ecosystems
abstract
Change introduces conflict into software ecosystems: breaking changes may ripple through the ecosystem and trigger rework for users of a package, but often developers can invest additional effort or accept opportunity costs to alleviate or delay downstream costs. We performed a multiple case study of three software ecosystems with different tooling and philosophies toward change, Eclipse, R/CRAN, and Node.js/npm, to understand how developers make decisions about change and change-related costs and what practices, tooling, and policies are used. We found that all three ecosystems differ substantially in their practices and expectations toward change and that those differences can be explained largely by different community values in each ecosystem. Our results illustrate that there is a large design space in how to build an ecosystem, its policies and its supporting infrastructure; and there is value in making community values and accepted tradeoffs explicit and transparent in order to resolve conflicts and negotiate change-related costs.
Christopher Bogart, Christian Kästner, James D. Herbsleb, Ferdian Thung
SIGSOFT FSE1
2015 Programs for people: What we can learn from lab protocols
abstract
Humans play an active role in the execution of certain kinds of programs, such as spreadsheets, workflows and interactive notebooks. Interacting closely with execution is especially useful when end-users are learning from examples while doing their work. In order to better understand the language features needed to support this kind of use, we investigated a particularly rigid and formalized category of “program” people write for each other: lab protocols. These protocols present a linear, idealized process despite the complex contingencies of the lab work they describe. However, they employ a variety of techniques for limiting or expanding the semantic interpretation of individual steps and for integrating outside protocols. We use these observations to derive implications for the design of interactive and mixed-initiative programming languages.
Keeley Abbott, Christopher Bogart, Eric Walkingshaw
VL/HCC2
2013 How Programmers Debug, Revisited: An Information Foraging Theory Perspective
abstract
Many theories of human debugging rely on complex mental constructs that offer little practical advice to builders of software engineering tools. Although hypotheses are important in debugging, a theory of navigation adds more practical value to our understanding of how programmers debug. Therefore, in this paper, we reconsider how people go about debugging in large collections of source code using a modern programming environment. We present an information foraging theory of debugging that treats programmer navigation during debugging as being analogous to a predator following scent to find prey in the wild. The theory proposes that constructs of scent and topology provide enough information to describe and predict programmer navigation during debugging, without reference to mental states such as hypotheses. We investigate the scope of our theory through an empirical study of 10 professional programmers debugging a real-world open source program. We found that the programmers' verbalizations far more often concerned scent-following than hypotheses. To evaluate the predictiveness of our theory, we created an executable model that predicted programmer navigation behavior more accurately than comparable models that did not consider information scent. Finally, we discuss the implications of our results for enhancing software engineering tools.
Joseph Lawrance, Christopher Bogart, Margaret M. Burnett, Rachel K. E. Bellamy, Kyle Rector, Scott D. Fleming
IEEE Trans. Software Eng.2
2012 Designing a debugging interaction language for cognitive modelers: an initial case study in natural programming plus
abstract
In this paper, we investigate how a debugging environment should support a population doing work at the core of HCI research: cognitive modelers. In conducting this investigation, we extended the Natural Programming methodology (a user-centered design method for HCI researchers of programming environments), to add an explicit method for mapping the outcomes of NP's empirical investigations to a language design. This provided us with a concrete way to make the design leap from empirical assessment of users' needs to a language. The contributions of our work are therefore: (1) empirical evidence about the content and sequence of cognitive modelers' information needs when debugging, (2) a new, empirically derived, design specification for a debugging interaction language for cognitive modelers, and (3) an initial case study of our "Natural Programming Plus" methodology.
Christopher Bogart, Margaret M. Burnett, Scott Douglass, Hannah Adams, Rachel White
CHI1
2012 Reactive information foraging: an empirical investigation of theory-based recommender systems for programmers
abstract
Information Foraging Theory (IFT) has established itself as an important theory to explain how people seek information, but most work has focused more on the theory itself than on how best to apply it. In this paper, we investigate how to apply a reactive variant of IFT (Reactive IFT) to design IFT-based tools, with a special focus on such tools for ill-structured problems. Toward this end, we designed and implemented a variety of recommender algorithms to empirically investigate how to help people with the ill-structured problem of finding where to look for information while debugging source code. We varied the algorithms based on scent type supported (words alone vs. words + code structure), and based on use of foraging momentum to estimate rapidity of foragers' goal changes. Our empirical results showed that (1) using both words and code structure significantly improved the ability of the algorithms to recommend where software developers should look for information; (2) participants used recommendations to discover new places in the code and also as shortcuts to navigate to known places; and (3) low-momentum recommendations were significantly more useful than high-momentum recommendations, suggesting rapid and numerous goal changes in this type of setting. Overall, our contributions include two new recommendation algorithms, empirical evidence about when and why participants found IFT-based recommendations useful, and implications for the design of tools based on Reactive IFT.
David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Bonnie E. John, Rachel K. E. Bellamy, Calvin Swart
CHI4
2011 Modeling programmer navigation: A head-to-head empirical evaluation of predictive models
abstract
Software developers frequently need to perform code maintenance tasks, but doing so requires time-consuming navigation through code. A variety of tools are aimed at easing this navigation by using models to identify places in the code that a developer might want to visit, and then providing shortcuts so that the developer can quickly navigate to those locations. To date, however, only a few of these models have been compared head-to-head to assess their predictive accuracy. In particular, we do not know which models are most accurate overall, which are accurate only in certain circumstances, and whether combining models could enhance accuracy. Therefore, we have conducted an empirical study to evaluate the accuracy of a broad range of models for predicting many different kinds of code navigations in sample maintenance tasks. Overall, we found that models tended to perform best if they took into account how recently a developer has viewed pieces of the code, and if models took into account the spatial proximity of methods within the code. We also found that the accuracy of single-factor models can be improved by combining factors, using a spreading-activation based approach, to produce multi-factor models. Based on these results, we offer concrete guidance about how these models could be used to provide enhanced software development tools that ease the difficulty of navigating through code.
David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Liza John, Christopher Bogart, Bonnie E. John, Margaret M. Burnett, Rachel K. E. Bellamy
VL/HCC5
2010 Reactive information foraging for evolving goals
abstract
Information foraging models have predicted the navigation paths of people browsing the web and (more recently) of programmers while debugging, but these models do not explicitly model users' goals evolving over time. We present a new information foraging model called PFIS2 that does model information seeking with potentially evolving goals. We then evaluated variants of this model in a field study that analyzed programmers' daily navigations over a seven-month period. Our results were that PFIS2 predicted users' navigation remarkably well, even though the goals of navigation, and even the information landscape itself, were changing markedly during the pursuit of information.
Joseph Lawrance, Margaret M. Burnett, Rachel K. E. Bellamy, Christopher Bogart, Calvin Swart
CHI4
2010 Debugging with Evaluation Abstractions
abstract
Do programmers put everything they know about a problem domain into the code they write to solve that problem? Or do they instead select a relatively parsimonious subset of that knowledge to define formally and translate into code? If the latter is true, it may be that they bring their original, richer perspective to bear in checking their program's behavior against their expectations. I will refer to these richer abstractions as "evaluation abstractions", in contrast to the formal "programming abstractions" embodied in the code.
Christopher Bogart
VL/HCC1
2010 Does My Model Work? Evaluation Abstractions of Cognitive Modelers
abstract
Are the abstractions that scientific modelers use to build their models in a modeling language the same abstractions they use to evaluate the correctness of their models? The extent to which such differences exist seems likely to correspond to additional effort of modelers in determining whether their models work as intended. In this paper, we therefore investigate the distinction between "programming abstractions" and "evaluation abstractions". As the basis of our investigation, we conducted a case study on cognitive modeling. We report modelers' evaluation abstractions, and the lengths they went to in evaluating their models. From these results, we derive design implications for several categories of persistent, first-class evaluation abstractions in future debugging tools for modelers.
Christopher Bogart, Margaret M. Burnett, Scott Douglass, David Piorkowski, Amber Shinsel
VL/HCC1
2009 Predicting reuse of end-user web macro scripts
abstract
Repositories of code written by end-user programmers are beginning to emerge, but when a piece of code is new or nobody has yet reused it, then current repositories provide users with no information about whether that code might be appropriate for reuse. Addressing this problem requires predicting reusability based on information that exists when a script is created. To provide such a model for web macro scripts, we identified script traits that might plausibly predict reuse, then used IBM CoScripter repository logs to statistically test how well each corresponded to reuse. We then built a machine learning model that combines the useful traits and evaluated how well it can predict four different types of reuse that we saw in the repository logs. Our model was able to predict reuse from a surprisingly small set of traits. It is simple enough to be explained in only 6-11 rules, making it potentially viable for integration in repository search engines for end-user programmers.
Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Allen Cypher, Brad A. Myers, Mary Shaw
VL/HCC2
2008 Rhetorical end-user programming
abstract
The study of rhetoric has long been a route to empowerment for people, by helping them share their ideas and inspire cooperation from others. This research discusses the possibility of designers of programming environments to piggyback on that success which would allow people without computer science training to leverage their human communication skills more fully when programming.
Christopher Bogart
VL/HCC1
2008 End-user programming in the wild: A field study of CoScripter scripts
abstract
Although a new class of languages has emerged to enable end users to create their own Web applications, little is known about how end-user programmers actually use such languages in the real world. In this paper, we report a field study on over 1400 scripts collected from the Internet which were created by early adopters of CoScripter, a Web macro programming-by-demonstration language. We contrast these Internet scripts with those written by users inside IBM, and describe script usage and re-usage patterns, features used, and users' clever workarounds for features not present in the language. The results show how users grapple with such programming notions as repetition, generalization, and reuse, sometimes inventing their own devices for these. Finally, we discuss the many scripts we found with social implications, whose purposes were to circumvent intended rules, regulations, and usage norm assumptions of a number of Web sites.
Christopher Bogart, Margaret M. Burnett, Allen Cypher, Christopher Scaffidi
VL/HCC1
2008 Can feature design reduce the gender gap in end-user software development environments?
abstract
Recent research has begun to report that female end-user programmers are often more reluctant than males to employ features that are useful for testing and debugging. These earlier findings suggest that, unless such features can be changed in some appropriate way, there are likely to be important gender differences in end-user programmerspsila benefits from these features. In this paper, we compare end-user programmerspsila feature usage in an environment that supports end-user debugging, against an extension of the same environment with two features designed to help ameliorate the effects of low self-efficacy. Our results show ways in which these features affect female versus male enduser programmerspsila self-efficacy, attitudes, usage of testing and debugging features, and performance.
Valentina Grigoreanu, Jill Cao, Todd Kulesza, Christopher Bogart, Kyle Rector, Margaret M. Burnett, Susan Wiedenbeck
VL/HCC4
1990 Genetic algorithms and neural networks: optimizing connections and connectivity
L. Darrell Whitley, Timothy Starkweather, Christopher Bogart
Parallel Comput.3