VLDB 2026 Research / reviewers in the wild / expert
Philip J. Guo
dblp:32/193 · also Philip Jia Guo
· DBLP profile ↗
80ranked-venue papers
17as first author
22since 2021 · last 2026
0000-0002-4579-5754ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 53 · 8 first-author · 19 since 2021Systems, architecture and hardware · 13 · 6 first-author · 2 since 2021Software engineering, systems software and programming languages · 13 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Desk: Barriers and Future Opportunities for AI to Assist Scientists in Embodied Physical TasksabstractMore scientists are now using AI, but prior studies have examined only how they use it ‘at the desk’ for computer-based work. However, given that scientific work often happens ‘beyond the desk’ at lab and field sites, we conducted the first study of how scientific practitioners use AI for embodied physical tasks. We interviewed 12 scientific practitioners doing hands-on lab and fieldwork in domains like nuclear fusion, primate cognition, and biochemistry, and found three barriers to AI adoption in these settings: 1) experimental setups are too high-stakes to risk AI errors, 2) constrained environments make it hard to use AI, and 3) AI cannot match the tacit knowledge of humans. Participants then developed speculative designs for future AI assistants to 1) monitor task status, 2) organize lab-wide knowledge, 3) monitor scientists’ health, 4) do field scouting, 5) do hands-on chores. Our findings point toward AI as background infrastructure to support physical work rather than replacing human expertise. Irene Hou, Alexander Qin, Lauren Cheng, Philip J. Guo |
CHI | 4 |
| 2026 | "Bespoke Bots": Diverse Instructor Needs for Customizing Generative AI Classroom ChatbotsabstractInstructors are increasingly experimenting with AI chatbots for classroom support. To investigate how instructors adapt chatbots to their own contexts, we first analyzed existing resources that provide prompts for educational purposes. We identified ten common categories of customization, such as persona, guardrails, and personalization. We then conducted interviews with ten university STEM instructors and asked them to card-sort the categories into priorities. We found that instructors consistently prioritized the ability to customize chatbot behavior to align with course materials and pedagogical strategies and de-prioritized customizing persona/tone. However, their prioritization of other categories varied significantly by course size, discipline, and teaching style, even across courses taught by the same individual, highlighting that no single design can meet all contexts. These findings suggest that modular AI chatbots may provide a promising path forward. We offer design implications for educational developers building the next generation of customizable classroom AI systems. Irene Hou, Zeyu Xiong, Philip J. Guo, April Yi Wang |
CHI | 3 |
| 2026 | Behind the Scenes of Delivering a Large Computing Course: The Experience of a TA Managing LogisticsabstractThere are many tasks that must be done to keep a computing course running smoothly. In addition to pedagogical work, instructors do a multitude of behind-the-scenes administrative and organizational work themselves or delegate it to Teaching Assistants (TAs). Despite the pervasiveness of this type of work, the explicit details of what kinds of tasks are necessary, the processes to execute these tasks well, and the experiences and feelings of those doing the tasks have not been reported in the literature. Therefore, this experience report surfaces these details by presenting the reflections of a long-time TA who has worked on significant administrative tasks across eight CS courses at our institution over the years. As a representative sample of her experiences, we report on three anecdotes that highlight logistical and internal challenges she faced as a TA for a CS1 course: 1) managing the flow of 600 students coming into a computer lab to take proctored assessments, 2) scanning and uploading thousands of pages of on-paper exams, 3) leading an exam grading session of dozens of TAs. Based on her reflections, we see how much intentional foresight and preparation goes into preventing chaotic situations, as well as many concurrent and overlapping threads to manage in an overarching task. The combination of these factors as well as challenges in team and interpersonal dynamics were the primary sources of stress. Our goal is to bring awareness to this aspect of course delivery that is oftentimes hidden and inspire future research around it. Rachel S. Lim, Philip J. Guo |
SIGCSE (1) | 2 |
| 2025 | Undergraduate Computing Tutors' Perceptions of their Roles, Stressors, and Barriers to EffectivenessabstractUndergraduate teaching assistants (tutors) are commonly employed in computing courses to help students with programming assignments. Prior research in computing education has reported the benefits of tutoring both for students and for the tutors' own learning. In contrast, recent research that examined actual tutoring sessions has reported that these sessions may be less productive than one might hope, with tutors often just giving students the answers to their problems without trying to teach the underlying concepts. To better understand why tutors may be employing these suboptimal practices, we interviewed ten tutors across early computing courses in higher education to identify their perceived role in these sessions, what stressors and factors influence their ability to perform their job effectively, and what kinds of best practices they learned in their tutor training course. Tutors reported their roles around student learning, gauging student understanding, identifying or providing solutions to students, and providing socioemotional support. They reported their stressors around environmental factors (e.g., number of students waiting to be helped, preparation time, peer-tutor frustrations), internal influences, student behavior, student skill levels, and feeling the need to ''read a student's mind.'' Regarding their tutor training course, Tutors reported learning about interaction guidelines and procedures and question-based problem solving. We conclude by discussing how these results may contribute to the less-effective behaviors seen in prior research and potential ways to improve tutoring in computing courses. Ismael Villegas Molina, Jeannie Kim, Audria Montalvo, Apollo Larragoitia, Rachel S. Lim, Philip J. Guo, Sophia Krause-Levy, Leo Porter 0001 |
SIGCSE (1) | 6 |
| 2025 | The Design Space of LLM-Based AI Coding Assistants: An Analysis of 90 Systems in Academia and IndustryabstractOver the past few years, millions of people have been using LLM-based AI tools to aid in programming, data analysis, and software engineering tasks. These AI coding assistants range from specialized tools like GitHub Copilot to general-purpose chatbots like Claude. In parallel, academics have published dozens of papers on forward-looking prototypes to expand our collective thinking beyond present-day industry trends. However, despite rapid advances in both sectors in recent years, we still lack an understanding of how their designs relate to one another and what tradeoffs are commonly made. At this key moment in 2025 when design patterns are starting to emerge, it is important to zoom out to see the forest instead of the trees. To do so, we performed the first comprehensive design analysis of 90 LLM-based AI coding assistants. We categorized the feature sets of 58 industry products and 32 academic projects, then formulated a design space that captures key variations in their user experiences. Our design space covers $\mathbf{1 0}$ dimensions related to UI modalities, system inputs, capabilities, and outputs. We use this design space to reveal trends in both industry and academic projects across three eras ranging from autocomplete to chat to agent-based interfaces. Lastly, to address the question of who the target users of these tools are, we present six user personas whose preferences lie in different regions of our design space: professional software engineers, HCI researchers and hobbyist programmers, UX designers, conversational programmers (e.g., product managers and marketers), data scientists, and students. Sam Lau, Philip J. Guo |
VL/HCC | 2 |
| 2024 | Taking ASCII Drawings Seriously: How Programmers Diagram CodeabstractDocumentation in codebases facilitates knowledge transfer. But tools for programming are largely text-based, and so developers resort to creating ASCII diagrams—graphical artifacts approximated with text—to show visual ideas within their code. Despite real-world use, little is known about these diagrams. We interviewed nine authors of ASCII diagrams, learning why they use ASCII and what roles the diagrams play. We also compile and analyze a corpus of 507 ASCII diagrams from four open source projects, deriving a design space with seven dimensions that classify what these diagrams show, how they show it, and ways they connect to code. These investigations reveal that ASCII diagrams are professional artifacts used across many steps in the development lifecycle, diverse in role and content, and used because they visualize ideas within the variety of programming tools in use. Our findings highlight the importance of visualization within code and lay a foundation for future programming tools that tightly couple text and graphics. Devamardeep Hayatpur, Brian Hempel, Kathy Chen, William Duan, Philip J. Guo, Haijun Xia |
CHI | 5 |
| 2024 | Perpetual Teaching Across Temporary Places: Conditions, Motivations, and Practices of Media Artists Teaching Computing WorkshopsabstractWhy and how do new media artists teach computing? Over the past decade, computing has become a part of the standard curriculum in university art and design departments, along with the advent of influential informal learning communities and self-organized schools. This paper is the first systematic attempt to map the diverse conditions, motivations, and practices of new media artists teaching computing. Interviews with 18 new media artists from 5 countries and 17 different sites revealed that teaching computing is closely integrated with their art practice, with a shared aim to cultivate new cultures in computing rather than only to transfer knowledge. We gathered new media artists’ accounts of precarious work, lack of time and place for their practices, and unrealistic expectations for instant results they face in their teaching. Within these precarious conditions, they developed a unique set of practices for “perpetual teaching,” which promotes self-reflective, critical, and situated learning. Our findings from this study are a call for further investigation of educators’ roles in creating cultures in computing, especially incorporating practices outside of conventional computing education settings. Alice Mira Chung, Philip J. Guo |
ICER (1) | 2 |
| 2024 | UNFOLD: Enabling Live Programming for Debugging GUI ApplicationsabstractDebugging GUI applications is challenging because of the difficult-to-debug state changes caused by asynchronous event handling and user interactions. Live programming, a paradigm where programmers continuously see real-time traces of every execution step as they edit the code, is promising for this context, as it automates the visualization of state changes upon inputs to the program. This paper explores how to design a live programming experience for debugging GUI applications and studies the effects of live programming in this context. Through a formative design exploration, we derive three core concepts for enabling live programming in debugging GUI applications: a UI states timeline, connections between the UI and the code, and automated event recording. We implemented these concepts in UNFOLD, a live programming environment for JavaScript-based GUI applications. A within-subject study with 12 participants shows that, with UNFOLD, participants locate bugs faster in tasks amenable to live programming, leverage liveness when debugging, and deem the tool helpful and easy to use. Ruanqianqian (Lisa) Huang, Philip J. Guo, Sorin Lerner |
VL/HCC | 2 |
| 2023 | From "Ban It Till We Understand It" to "Resistance is Futile": How University Programming Instructors Plan to Adapt as More Students Use AI Code Generation and Explanation Tools such as ChatGPT and GitHub CopilotabstractOver the past year (2022–2023), recently-released AI tools such as ChatGPT and GitHub Copilot have gained significant attention from computing educators. Both researchers and practitioners have discovered that these tools can generate correct solutions to a variety of introductory programming assignments and accurately explain the contents of code. Given their current capabilities and likely advances in the coming years, how do university instructors plan to adapt their courses to ensure that students still learn well? To gather a diverse sample of perspectives, we interviewed 20 introductory programming instructors (9 women + 11 men) across 9 countries (Australia, Botswana, Canada, Chile, China, Rwanda, Spain, Switzerland, United States) spanning all 6 populated continents. To our knowledge, this is the first empirical study to gather instructor perspectives about how they plan to adapt to these AI coding tools that more students will likely have access to in the future. We found that, in the short-term, many planned to take immediate measures to discourage AI-assisted cheating. Then opinions diverged about how to work with AI coding tools longer-term, with one side wanting to ban them and continue teaching programming fundamentals, and the other side wanting to integrate them into courses to prepare students for future jobs. Our study findings capture a rare snapshot in time in early 2023 as computing instructors are just starting to form opinions about this fast-growing phenomenon but have not yet converged to any consensus about best practices. Using these findings as inspiration, we synthesized a diverse set of open research questions regarding how to develop, deploy, and evaluate AI coding tools for computing education. Sam Lau, Philip J. Guo |
ICER (1) | 2 |
| 2023 | Uncovering the Hidden Curriculum of University Computing Majors via Undergraduate-Written Mentoring Guides: A Learner-Centered Design WorkflowabstractThe hidden curriculum consists of the unwritten rules, unspoken norms, and field-specific insider knowledge that are essential for student success but are not taught in classes. Examples include social norms about how to interact with authority figures, where to ask for unadvertised career-related opportunities, and how to navigate around the official rules of a bureaucracy. The hidden curriculum can be pervasive in university computing majors because some students come in with more prior childhood exposure to technology culture and can thus navigate this cultural context more fluently. It is possible to learn this type of tacit knowledge from personal mentors, but not everyone has access to a good mentor. To address this challenge, this paper presents a novel thesis for how to teach students the hidden curriculum in a more scalable way: We propose that a peer-written guide that has a relatable tone and a focus on local context can emulate what a peer mentor does by emotionally resonating with students, teaching them aspects of the hidden curriculum, and motivating them to take concrete action. To demonstrate this thesis we created a mentoring guide for interdisciplinary computing HCI majors at our university. Interviews with 17 students and a survey of 112 students showed that our guide’s relatable tone could emotionally resonate with students, that it boosted some readers’ self-confidence, and that it inspired them to take actions such as creating a project portfolio. Based on these experiences, we developed a five-step learner-centered design workflow to help others create guides for their own local contexts, with recommendations for 1) setting up a mentoring guide, 2) needfinding, 3) creating, 4) distributing, and 5) maintaining. Kendall Nakai, Philip J. Guo |
ICER (1) | 2 |
| 2023 | ForewordabstractWelcome to the 2023 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC). We proudly continue VL/HCC's tradition as the premier international forum for research on how people learn, express, and understand computational ideas and on the languages, tools, and interventions that can aid people in doing so. Our program this year was diverse as usual, and we also encouraged submissions and keynote talks around low-code / no-code programming systems, especially given the recent popularity of AI coding assistance tools based on large language models (LLMs) such as GitHub Copilot, ChatGPT, and others. With these new technologies developing at such a rapid pace, it is an exciting time to be learning and doing programming. Thomas D. LaToza, Esther Guerra, Philip J. Guo |
VL/HCC | 3 |
| 2022 | The Design Space of Livestreaming Equipment Setups: Tradeoffs, Challenges, and OpportunitiesabstractLivestreaming has grown popular in recent years, with millions of people broadcasting themselves making digital art, playing games, programming, and doing other activities on sites like Twitch and YouTube. While many researchers have studied the actions of both streamers and their viewers, to our knowledge there has been no comprehensive analysis of the actual hardware and software equipment used in livestreaming. In this survey paper we present a holistic overview of modern livestreaming equipment in 2022 by analyzing 40 videos where streamers talk about various aspects of their setups. We categorized their equipment choices into a design space with ten dimensions: computer, software, stream control, encoding, cameras, lighting, video accessories, microphones, audio mixers, and audio accessories. We found that each streamer must make tradeoffs between lower- and higher-fidelity options within each dimension. Our design space analysis can inform ideas for future streaming support tools and, more broadly, tools for remote collaboration and learning via live video. As more of us work and learn online, we are in essence becoming amateur livestreamers, so understanding how professional streamers use their equipment to effectively engage their audiences might help us also engage better with our coworkers and classmates. Ian Drosos, Philip J. Guo |
Conference on Designing Interactive Systems | 2 |
| 2022 | The Challenges of Evolving Technical Courses at Scale: Four Case Studies of Updating Large Data Science CoursesabstractInstructors who teach large-scale technical courses, especially on data science and programming, must do a large amount of logistical work when updating their courses. All of this behind-the-scenes labor takes time away from the pedagogically-meaningful work of teaching students. Over the past five years, the authors of this paper have created and updated eight courses for an undergraduate data science program that serves over 2,000 students per year. We present four case studies from our teaching experiences that highlight major challenges in maintaining and updating technical courses: 1) There were intricate dependencies between course materials, so making updates to one part of the course would require updating many other parts. 2) We needed to maintain several variants of course materials such as assignments. 3) We wrote large amounts of ad-hoc custom software infrastructure to manage logistics. 4) We could not easily reuse software written by others. Our case studies point to design ideas for instructor-oriented tools that can reduce the logistical complexities of teaching at scale, thus letting instructors focus on the substance of teaching rather than on mundane logistics. Sam Lau, Justin Eldridge, Shannon Ellis, Aaron Fraenkel, Marina Langlois, Suraj Rampure, Janine Tiefenbruck, Philip J. Guo |
L@S | 8 |
| 2022 | Scaling Up Access to the Hidden Curriculum: A Design Methodology for Creating Undergraduate Mentoring GuidesabstractThe hidden curriculum consists of the unwritten rules, unspoken norms, and field-specific insider knowledge that are essential for student success but are not taught in classes. Examples include how to approach professors and prospective employers to ask for opportunities and how to gather information that is relevant to one's career goals. Students now informally learn these skills from more experienced peers, but not everyone has access to personalized one-on-one mentoring. We scaled up access to the hidden curriculum by creating a mentoring guide to advise students majoring in HCI/Design at our university. Readers have found it useful for orienting new students, helping older students who feel behind, and serving as a confidence booster. We synthesized our experiences into a five-step design methodology to help other students to create peer mentoring guides, with recommendations for 1) setup, 2) needfinding, 3) creating, 4) distributing, and 5) maintaining. Kendall Nakai, Philip J. Guo |
L@S | 2 |
| 2022 | Five Pedagogical Principles of a User-Centered Design Course that Prepares Computing Undergraduates for Industry JobsabstractWe present a new user-centered design course that prepares computing undergraduates for software industry jobs such as UI/UX designer, product designer, and product manager. Our course aims to bridge the academia-industry gap and innovates upon prior published HCI courses due to its targeted focus on job preparation, inclusion, and scale. Nearly 200 students (55% women) have taken it in the past two years. We developed its curriculum to align with the needs of modern industry employers and implemented five theory-backed pedagogical principles: 1) industry-relevant project prompts developed in consultation with recent course alumni, 2) final project deliverable optimized for job-seeking, 3) no coding required to foster inclusion, 4) low-stress effort-based grading to further foster inclusion, 5) weekly feedback and chances for revisions. We discuss the theoretical rationale behind these five principles and how instructors can potentially apply them to a broad range of project-based courses across many areas of computing. Sean Kross, Philip J. Guo |
SIGCSE (1) | 2 |
| 2022 | How Computer Science and Statistics Instructors Approach Data Science Pedagogy Differently: Three Case StudiesabstractOver the past decade, data science courses have been growing more popular across university campuses. These courses often involve a mix of programming and statistics and are taught by instructors from diverse backgrounds. In our experiences launching a data science program at a large public U.S. university over the past four years, we noticed one central tension within many such courses: instructors must finely balance how much computing versus statistics to teach in the limited available time. In this experience report, we provide a detailed firsthand reflection on how we have personally balanced these two major topic areas within several offerings of a large introductory data science course that we taught and wrote an accompanying textbook for; our course has served several thousand students over the past four years. We present three case studies from our experiences to illustrate how computer science and statistics instructors approach data science differently on topics ranging from algorithmic depth to modeling to data acquisition. We then draw connections to deeper tradeoffs in data science to help guide instructors who design interdisciplinary courses. We conclude by suggesting ways that instructors can incorporate both computer science and statistics perspectives to improve data science teaching. Sam Lau, Deborah Nolan, Joseph Gonzalez 0001, Philip J. Guo |
SIGCSE (1) | 4 |
| 2022 | "There's no way to keep up!": Diverse Motivations and Challenges Faced by Informal Learners of MLabstractIn recent years, more people from different backgrounds are trying to informally learn Machine Learning (ML) using a plethora of online resources, yet we know little about their motivations and learning strategies. We carried out interviews with 22 informal learners of ML from diverse job roles and backgrounds, including Computer Science, Medicine, Finance, and others, to understand their approaches, preferences, and challenges in locating and interacting with different resources to manage their learning. We analyzed our findings using the framework of self-directed learning and found that these informal learners struggled in all stages of self-direction, including identifying learning goals and selecting resources, and that their challenges were most acute in the last stage of gauging progress and evaluating outcomes. We identify several opportunities for future research to better understand and support informal learners of ML (and other complex technical skills). In particular, there is a need to foster more self-monitoring and self-reflection techniques that can help informal learners become more self-aware and effective in directing their learning. Rimika Chaudhury, Philip J. Guo, Parmit K. Chilana |
VL/HCC | 2 |
| 2022 | Seq2Parse: neurosymbolic parse error repairabstractWe present Seq2Parse, a language-agnostic neurosymbolic approach to automatically repairing parse errors. Seq2Parse is based on the insight that Symbolic Error Correcting (EC) Parsers can, in principle, synthesize repairs, but, in practice, are overwhelmed by the many error-correction rules that are not relevant to the particular program that requires repair. In contrast, Neural approaches are fooled by the large space of possible sequence level edits, but can precisely pinpoint the set of EC-rules that are relevant to a particular program. We show how to combine their complementary strengths by using neural methods to train a sequence classifier that predicts the small set of relevant EC-rules for an ill-parsed program, after which, the symbolic EC-parsing algorithm can make short work of generating useful repairs. We train and evaluate our method on a dataset of 1,100,000 Python programs, and show that Seq2Parse is accurate and efficient : it can parse 94% of our tests within 2.1 seconds, while generating the exact user fix in 1 out 3 of the cases; and useful : humans perceive both Seq2Parse-generated error locations and repairs to be almost as good as human-generated ones in a statistically-significant manner. Georgios Sakkas, Madeline Endres, Philip J. Guo, Westley Weimer, Ranjit Jhala |
Proc. ACM Program. Lang. | 3 |
| 2021 | Inside the Mind of a CS Undergraduate TA: A Firsthand Account of Undergraduate Peer Tutoring in Computer LabsabstractAs CS enrollments continue to grow, introductory courses are employing more undergraduate TAs. One of their main roles is performing one-on-one tutoring in the computer lab to help students understand and debug their programming assignments. What goes on in the mind of an undergraduate TA when they are helping students with programming? In this experience report, we present firsthand accounts from an undergraduate TA documenting her 36 hours of in-lab tutoring for a CS2 course, where she engaged in 69 one-on-one help sessions. This report provides a unique perspective from an undergraduate's point-of-view rather than a faculty member's. We summarize her experiences by constructing a four-part model of tutoring interactions: a) The tutor begins the session with an initial state of mind (e.g., their energy/focus level, perceived time pressure). b) They observe the student's outward state upon arrival (e.g., how much they seem to care about learning). c) Using that observation, the tutor infers what might be going on inside the student's mind. d) The combination of what goes on inside the tutor's and student's minds affects tutoring interactions, which progress from diagnosis to planning to an explain-code-react loop to post-resolution activities. We conclude by discussing ways that this model can be used to design scaffolding for training novice TAs and software tools to help TAs scale their efforts to larger classes. Julia M. Markel, Philip J. Guo |
SIGCSE | 2 |
| 2021 | Ten Million Users and Ten Years Later: Python Tutor's Design Guidelines for Building Scalable and Sustainable Research Software in AcademiaabstractResearch software is often built as prototypes that never get widespread usage and are left unmaintained after a few papers get published. To counteract this trend, we propose a method for building research software with scale and sustainability in mind so that it can organically grow a large userbase and enable longer-term research. To illustrate this method, we present the design and implementation of Python Tutor (pythontutor.com), a code visualization tool that is, to our knowledge, one of the most widely-used pieces of research software developed within a university lab. Over the past decade, it has been used by over ten million people in over 180 countries. It has also contributed to 55 publications from 35 research groups in 13 countries. We distilled lessons from working on Python Tutor into three sets of design guidelines: 1) user experience design for scale and sustainability, 2) software architecture design for long-term sustainability, and 3) designing a sustainable software development workflow within academia. These guidelines can enable a student to create long-lasting software that reaches many users and facilitates research from many independent groups. Philip J. Guo |
UIST | 1 |
| 2021 | Streamers Teaching Programming, Art, and Gaming: Cognitive Apprenticeship, Serendipitous Teachable Moments, and Tacit Expert KnowledgeabstractLivestreaming is now a popular way for programmers, artists, and gamers to teach their craft online. In this paper we propose the idea that streaming can enable cognitive apprenticeship, a form of teaching where an expert works on authentic tasks while thinking aloud to explain their creative process. To understand how streamers teach in this naturalistic way, we performed a content analysis of 20 stream videos across four popular categories: web development, data science, digital art, and gaming. We discovered four kinds of serendipitous teachable moments that are reminiscent of cognitive apprenticeship: 1) creators encountered unexpected errors that led to improvised problem solving, 2) they generated improvised examples on-the-fly, 3) they sometimes went on insightful tangents, 4) they paused to give high-level advice that was contextualized within the work they were currently performing. We also found missed opportunities for additional teachable moments due to creators not being able to express their tacit (unspoken) expert knowledge because of pattern irreducibility, context dependence, and routinization. Ian Drosos, Philip J. Guo |
VL/HCC | 2 |
| 2021 | Orienting, Framing, Bridging, Magic, and Counseling: How Data Scientists Navigate the Outer Loop of Client Collaborations in Industry and AcademiaabstractData scientists often collaborate with clients to analyze data to meet a client's needs. What does the end-to-end workflow of a data scientist's collaboration with clients look like throughout the lifetime of a project? To investigate this question, we interviewed ten data scientists (5 female, 4 male, 1 non-binary) in diverse roles across industry and academia. We discovered that they work with clients in a six-stage outer-loop workflow, which involves 1) laying groundwork by building trust before a project begins, 2) orienting to the constraints of the client's environment, 3) collaboratively framing the problem, 4) bridging the gap between data science and domain expertise, 5) the inner loop of technical data analysis work, 6) counseling to help clients emotionally cope with analysis results. This novel outer-loop workflow contributes to CSCW by expanding the notion of what collaboration means in data science beyond the widely-known inner-loop technical workflow stages of acquiring, cleaning, analyzing, modeling, and visualizing data. We conclude by discussing the implications of our findings for data science education, parallels to design work, and unmet needs for tool development. Sean Kross, Philip J. Guo |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2020 | Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsabstractData wrangling is a difficult and time-consuming activity in computational notebooks, and existing wrangling tools do not fit the exploratory workflow for data scientists in these environments. We propose a unified interaction model based on programming-by-example that generates readable code for a variety of useful data transformations, implemented as a Jupyter notebook extension called Wrex. User study results demonstrate that data scientists are significantly more effective and efficient at data wrangling with Wrex over manual programming. Qualitative participant feedback indicates that Wrex was useful and reduced barriers in having to recall or look up the usage of various data transform functions. The synthesized code allowed data scientists to verify the intended data transformation, increased their trust and confidence in Wrex, and fit seamlessly within their cell-based notebook workflows. This work suggests that presenting readable code to professional data scientists is an indispensable component of offering data wrangling tools in notebooks. Ian Drosos, Titus Barik, Philip J. Guo, Robert DeLine, Sumit Gulwani |
CHI | 3 |
| 2020 | Learnersourcing at Scale to Overcome Expert Blind Spots for Introductory Programming: A Three-Year Deployment Study on the Python Tutor WebsiteabstractIt is hard for experts to create good instructional resources due to a phenomenon known as the expert blind spot: They forget what it was like to be a novice, so they cannot pinpoint exactly where novices commonly struggle and how to best phrase their explanations. To help overcome these expert blind spots for computer programming topics, we created a learnersourcing system that elicits explanations of misconceptions directly from learners while they are coding. We have deployed this system for the past three years to the widely-used Python Tutor coding website (pythontutor.com) and collected 16,791 learner-written explanations. To our knowledge, this is the largest dataset of explanations for programming misconceptions. By inspecting this dataset, we found surprising insights that we did not originally think of due to our own expert blind spots as programming instructors. We are now using these insights to improve compiler and run-time error messages to explain common novice misconceptions. Philip J. Guo, Julia M. Markel |
L@S | 1 |
| 2020 | The Design Space of Computational Notebooks: An Analysis of 60 Systems in Academia and IndustryabstractComputational notebooks such as Jupyter are now used by millions of data scientists, machine learning engineers, and computational researchers to do exploratory and end-user programming. In recent years, dozens of different notebook systems have been developed across academia and industry. However, we still lack an understanding of how their individual designs relate to one another and what their tradeoffs are. To provide a holistic view of this rapidly-emerging landscape, we performed, to our knowledge, the first comprehensive design analysis of dozens of notebook systems. We analyzed 60 notebooks (16 academic papers, 29 industry products, and 15 experimental/R&D projects) and formulated a design space that succinctly captures variations in system features. Our design space covers 10 dimensions that include diverse ways of importing data, editing code and prose, running code, and publishing notebook outputs. We conclude by suggesting ways for researchers to push future projects beyond the current bounds of this space. Sam Lau, Ian Drosos, Julia M. Markel, Philip J. Guo |
VL/HCC | 4 |
| 2019 | Practitioners Teaching Data Science in Industry and Academia: Expectations, Workflows, and ChallengesabstractData science has been growing in prominence across both academia and industry, but there is still little formal consensus about how to teach it. Many people who currently teach data science are practitioners such as computational researchers in academia or data scientists in industry. To understand how these practitioner-instructors pass their knowledge onto novices and how that contrasts with teaching more traditional forms of programming, we interviewed 20 data scientists who teach in settings ranging from small-group workshops to large online courses. We found that: 1) they must empathize with a diverse array of student backgrounds and expectations, 2) they teach technical workflows that integrate authentic practices surrounding code, data, and communication, 3) they face challenges involving authenticity versus abstraction in software setup, finding and curating pedagogically-relevant datasets, and acclimating students to live with uncertainty in data analysis. These findings can point the way toward better tools for data science education and help bring data literacy to more people around the world. Sean Kross, Philip J. Guo |
CHI | 2 |
| 2019 | Theory and practice of string solvers (invited talk abstract)abstractThe paper titled "Hampi: A Solver for String Constraints" was published in the proceedings of the International Symposium on Software Testing and Analysis (ISSTA) 2009, and has been selected to receive the ISSTA 2019 Impact Paper Award. The paper describes HAMPI, one of the first practical solver aimed at solving the satisfiability problem for a theory of string (word) equations, operations over strings, predicates over regular expressions and context-free grammars. HAMPI has been used widely to solve many software engineering and security problems, and has inspired considerable research on string solving algorithms and their applications. Adam Kiezun, Philip J. Guo, Pieter Hooimeijer, Michael D. Ernst, Vijay Ganesh 0001 |
ISSTA | 2 |
| 2019 | Improv: Teaching Programming at Scale via Live CodingabstractComputer programming instructors frequently perform live coding in settings ranging from MOOC lecture videos to online livestreams. However, there is little tool support for this mode of teaching, so presenters must now either screen-share or use generic slideshow software. To overcome the limitations of these formats, we propose that programming environments should directly facilitate live coding for education. We prototyped this idea by creating Improv, an IDE extension for preparing and delivering code-based presentations informed by Mayer's principles of multimedia learning. Improv lets instructors synchronize blocks of code and output with slides and create preset waypoints to guide their presentations. A case study on 30 educational videos containing 28 hours of live coding showed that Improv was versatile enough to replicate approximately 96% of the content within those videos. In addition, a preliminary user study on four teaching assistants showed that Improv was expressive enough to allow them to make their own custom presentations in a variety of styles and improvise by live coding in response to simulated audience questions. Users mentioned that Improv lowered cognitive load by minimizing context switching and made it easier to fix errors on-the-fly than using slide-based presentations. Charles H. Chen, Philip J. Guo |
L@S | 2 |
| 2019 | Bespoke: Interactively Synthesizing Custom GUIs from Command-Line Applications By DemonstrationabstractProgrammers, researchers, system administrators, and data scientists often build complex workflows based on command-line applications. To give these power users the well-known benefits of GUIs, we created Bespoke, a system that synthesizes custom GUIs by observing user demonstrations of command-line apps. Bespoke unifies the two main forms of desktop human-computer interaction (command-line and GUI) via a hybrid approach that combines the flexibility and composability of the command line with the usability and discoverability of GUIs. To assess the versatility of Bespoke, we ran an open-ended study where participants used it to create their own GUIs in domains that personally motivated them. They made a diverse set of GUIs for use cases such as cloud computing management, machine learning prototyping, lecture video transcription, integrated circuit design, remote code deployment, and gaming server management. Participants reported that the benefit of these bespoke GUIs was that they exposed only the most relevant subset of options required for their specific needs. In contrast, vendor-made GUIs usually include far more panes, menus, and settings since they must accommodate a wider range of use cases. Priyan Vaithilingam, Philip J. Guo |
UIST | 2 |
| 2019 | Mallard: Turn the Web into a Contextualized Prototyping Environment for Machine LearningabstractMachine learning (ML) can be hard to master, but what first trips up novices is something much more mundane: the incidental complexities of installing and configuring software development environments. Everyone has a web browser, so can we let people experiment with ML within the context of any webpage they visit? This paper's contribution is the idea that the web can serve as a contextualized prototyping environment for ML by enabling analyses to occur within the context of data on actual webpages rather than in isolated silos. We realized this idea by building Mallard, a browser extension that scaffolds acquiring and parsing web data, prototyping with pretrained ML models, and augmenting webpages with ML-driven results and interactions. To demonstrate the versatility of Mallard, we performed a case study where we used it to prototype nine ML-based browser apps, including augmenting Amazon and Twitter websites with sentiment analysis, augmenting restaurant menu websites with OCR-based search, using real-time face tracking to control a Pac-Man game, and style transfer on Google image search results. These case studies show that Mallard is capable of supporting a diverse range of hobbyist-level ML prototyping projects. Philip J. Guo |
UIST | 2 |
| 2019 | Software Developers Learning Machine Learning: Motivations, Hurdles, and DesiresabstractThe growing popularity of machine learning (ML) has attracted more software developers to now want to adopt ML into their own practices, through tinkering with and learning from ML framework websites and online code examples. To investigate the motivations, hurdles, and desires of these software developers, we deployed a survey to the website of the TensorFlow.js ML framework. We found via 645 responses that many wanted to learn ML for aspirational reasons rather than for immediate job needs. Critically, developers faced hurdles due to a perceived lack of mathematical and theoretical background. They desired frameworks to provide more basic ML conceptual support, such as a curated corpus of best practices, conceptual tutorials, and a de-mystification of mathematical jargon into practical tips. These findings inform the design of ML frameworks and informal learning resources to broaden the base of people acquiring this increasingly important skill set. Carrie J. Cai, Philip J. Guo |
VL/HCC | 2 |
| 2019 | End-User Programmers Repurposing End-User Programming Tools to Foster Diversity in Adult End-User Programming EducationabstractEfforts to improve diversity in computing have mostly focused on K-12 and university student populations, so there is a lack of research on how to provide these benefits to adults who are not in school. To address this knowledge gap, we present a case study of how a nine-member team of end-user programmers designed an educational program to bring job-relevant computing skills to adult populations that have traditionally not been reached by existing efforts. This team conceived, implemented, and delivered Cloud Based Data Science (CBDS), a data science course designed for adults in their local community in historically marginalized groups that are underrepresented in computing fields. Notably, nobody on the course development team was a full-time educator or software engineer. To reduce the amount of time and cost required to launch their program, they repurposed end-user programming skills and tools from their professions, such as data-analytic programming and reproducible scientific research workflows. This case study demonstrates how the spirit of end-user programming can be a vehicle to drive social change through grassroots efforts. Sean Kross, Philip J. Guo |
VL/HCC | 2 |
| 2018 | Non-Native English Speakers Learning Computer Programming: Barriers, Desires, and Design OpportunitiesabstractPeople from nearly every country are now learning computer programming, yet the majority of programming languages, libraries, documentation, and instructional materials are in English. What barriers do non-native English speakers face when learning from English-based resources? What desires do they have for improving instructional materials? We investigate these questions by deploying a survey to a programming education website and analyzing 840 responses spanning 86 countries and 74 native languages. We found that non-native English speakers faced barriers with reading instructional materials, technical communication, reading and writing code, and simultaneously learning English and programming. They wanted instructional materials to use simplified English without culturally-specific slang, to use more visuals and multimedia, to use more culturally-agnostic code examples, and to embed inline dictionaries. Programming also motivated some to learn English better and helped clarify logical thinking about natural languages. Based on these findings, we recommend learner-centered design improvements to programming-related instructional resources and tools to make them more accessible to people around the world. Philip J. Guo |
CHI | 1 |
| 2018 | Mismatch of Expectations: How Modern Learning Resources Fail Conversational ProgrammersabstractConversational programmers represent a class of learners who are not required to write any code, yet try to learn programming to improve their participation in technical conversations. We carried out interviews with 23 conversational programmers to better understand the challenges they face in technical conversations, what resources they choose to learn programming, how they perceive the learning process, and to what extent learning programming actually helps them. Among our key findings, we found that conversational programmers often did not know where to even begin the learning process and ended up using formal and informal learning resources that focus largely on programming syntax and logic. However, since the end goal of conversational programmers was not to build artifacts, modern learning resources usually failed these learners in their pursuits of improving their technical conversations. Our findings point to design opportunities in HCI to invent learner-centered approaches that address the needs of conversational programmers and help them establish common ground in technical conversations. April Yi Wang, Ryan Mitts, Philip J. Guo, Parmit K. Chilana |
CHI | 3 |
| 2018 | Codemotion: expanding the design space of learner interactions with computer programming tutorial videosabstractLove them or hate them, videos are a pervasive format for delivering online education at scale. They are especially popular for computer programming tutorials since videos convey expert narration alongside the dynamic effects of editing and running code. However, these screencast videos simply consist of raw pixels, so there is no way to interact with the code embedded inside of them. To expand the design space of learner interactions with programming videos, we developed Codemotion, a computer vision algorithm that automatically extracts source code and dynamic edits from existing videos. Codemotion segments a video into regions that likely contain code, performs OCR on those segments, recognizes source code, and merges together related code edits into contiguous intervals. We used Codemotion to build a novel video player and then elicited interaction design ideas from potential users by running an elicitation study with 10 students followed by four participatory design workshops with 12 additional students. Participants collectively generated ideas for 28 kinds of interactions such as inline code editing, code-based skimming, pop-up video search, and in-video coding exercises. Kandarp Khandwala, Philip J. Guo |
L@S | 2 |
| 2018 | Students, systems, and interactions: synthesizing the first four years of learning@scale and charting the futureabstractWe survey all four years of papers published so far at the Learning at Scale conference in order to reflect on the major research areas that have been investigated and to chart possible directions for future study. We classified all 69 full papers so far into three categories: Systems for Learning at Scale, Interactions with Sociotechnical Systems, and Understanding Online Students. Systems papers presented technologies that varied by how much they amplify human effort (e.g., one-to-one, one-to-many, many-to-many). Interaction papers studied both individual and group interactions with learning technologies. Finally, student-centric study papers focused on modeling knowledge and on promoting global access and equity. We conclude by charting future research directions related to topics such as going beyond the MOOC hype cycle, axes of scale for systems, more immersive course experiences, learning on mobile devices, diversity in student personas, students as co-creators, and fostering better social connections amongst students. Sean Kross, Philip J. Guo |
L@S | 2 |
| 2018 | Porta: Profiling Software Tutorials Using Operating-System-Wide Activity TracingabstractIt can be hard for tutorial creators to get fine-grained feedback about how learners are actually stepping through their tutorials and which parts lead to the most struggle. To provide such feedback for technical software tutorials, we introduce the idea of tutorial profiling, which is inspired by software code profiling. We prototyped this idea in a system called Porta that automatically tracks how users navigate through a tutorial webpage and what actions they take on their computer such as running shell commands, invoking compilers, and logging into remote servers. Porta surfaces this trace data in the form of profiling visualizations that augment the tutorial with heatmaps of activity hotspots and markers that expand to show event details, error messages, and embedded screencast videos of user actions. We found through a user study of 3 tutorial creators and 12 students who followed their tutorials that Porta enabled both the tutorial creators and the students to provide more specific, targeted, and actionable feedback about how to improve these tutorials. Porta opens up possibilities for performing user testing of technical documentation in a more systematic and scalable way. Alok Mysore, Philip J. Guo |
UIST | 2 |
| 2018 | Fusion: Opportunistic Web Prototyping with UI MashupsabstractModern web development is rife with complexity at all layers, ranging from needing to configure backend services to grappling with frontend frameworks and dependencies. To lower these development barriers, we introduce a technique that enables people to prototype opportunistically by borrowing pieces of desired functionality from across the web without needing any access to their underlying codebases, build environments, or server backends. We implemented this technique in a browser extension called Fusion, which lets users create web UI mashups by extracting components from existing unmodified webpages and hooking them together using transclusion and JavaScript glue code. We demonstrate the generality and versatility of Fusion via a case study where we used it to create seven UI mashups in domains such as programming tools, data science, web design, and collaborative work. Our mashups include replicating portions of prior HCI systems (Blueprint for in-situ code search and DS.js for in-browser data science), extending the p5.js IDE for Processing with real-time collaborative editing, and integrating Python Tutor code visualizations into static tutorials. These UI mashups each took less than 15 lines of JavaScript glue code to create with Fusion. Philip J. Guo |
UIST | 2 |
| 2018 | The Impact of Culture on Learner Behavior in Visual DebuggersabstractPeople around the world are learning to code using online resources. However, research has found that these learners might not gain equal benefit from such resources, in particular because culture may affect how people learn from and use online resources. We therefore expect to see cultural differences in how people use and benefit from visual debuggers. We investigated the use of one popular online debugger which allows users to execute Python code and navigate bidirectionally through the execution using forward-steps and back-steps. We examined behavioral logs of 78,369 users from 69 countries and conducted an experiment with 522 participants from 82 countries. We found that people from countries that tend to prefer self-directed learning (such as those from countries with a low Power Distance, which tend to be less hierarchical than others) used about twice as many back-steps. We also found that for individuals whose values aligned with instructor-directed learning (those who scored high on a “Conservation” scale), back-steps were associated with less debugging success. Kyle Thayer, Philip J. Guo, Katharina Reinecke |
VL/HCC | 2 |
| 2017 | Older Adults Learning Computer Programming: Motivations, Frustrations, and Design OpportunitiesabstractComputer programming is a highly in-demand skill, but most learn-to-code initiatives and research target some of the youngest members of society: children and college students. We present the first known study of older adults learning computer programming. Using an online survey with 504 respondents aged 60 to 85 who are from 52 different countries, we discovered that older adults were motivated to learn to keep their brains challenged as they aged, to make up for missed opportunities during youth, to connect with younger family members, and to improve job prospects. They reported frustrations including a perceived decline in cognitive abilities, lack of opportunities to interact with tutors and peers, and trouble dealing with constantly-changing software technologies. Based on these findings, we propose a learner-centered design of techniques and tools for motivating older adults to learn programming and discuss broader societal implications of a future where more older adults have access to computer programming -- not merely computer literacy -- as a skill set. Philip J. Guo |
CHI | 1 |
| 2017 | CodePilot: Scaffolding End-to-End Collaborative Software Development for Novice ProgrammersabstractNovice programmers often have trouble installing, configuring, and managing disparate tools (e.g., version control systems, testing infrastructure, bug trackers) that are required to become productive in a modern collaborative software development environment. To lower the barriers to entry into software development, we created a prototype IDE for novices called CodePilot, which is, to our knowledge, the first attempt to integrate coding, testing, bug reporting, and version control management into a real-time collaborative system. CodePilot enables multiple users to connect to a web-based programming session and work together on several major phases of software development. An eight-subject exploratory user study found that first-time users of CodePilot spontaneously used it to assume roles such as developer/tester and developer/assistant when creating a web application together in pairs. Users felt that CodePilot could aid in scaffolding for novices, situational awareness, and lowering barriers to impromptu collaboration. Jeremy Warner, Philip J. Guo |
CHI | 2 |
| 2017 | Hack.edu: Examining How College Hackathons Are Perceived By Student Attendees and Non-AttendeesabstractCollege hackathons have become popular in the past decade, with tens of thousands of students now participating each year across hundreds of campuses. Since hackathons are informal learning environments where students learn and practice coding without any faculty supervision, they are an important site for computing education researchers to study as a complement to studying formal classroom learning environments. However, despite their popularity, little is known about why students choose to attend these events, what they gain from attending, and conversely, why others choose *not* to attend. This paper presents a mixed methods study that examines student perceptions of college hackathons by focusing on three main questions: 1.) Why are students motivated to attend hackathons? 2.) What kind of learning environment do these events provide? 3.) What factors discourage students from attending? Through semi-structured interviews with six college hackathon attendees (50% female), direct observation at a hackathon, and 256 survey responses from college students (42% female), we discovered that students were motivated to attend for both social and technical reasons, that the format generated excitement and focus, and that learning occurred incidentally, opportunistically, and from peers. Those who chose not to attend or had negative experiences cited discouraging factors such as physical discomfort, lack of substance, an overly competitive climate, an unwelcoming culture, and fears of not having enough prior experience. We conclude by discussing ideas for making college hackathons more broadly inclusive and welcoming in light of our study's findings. Jeremy Warner, Philip J. Guo |
ICER | 2 |
| 2017 | Omnicode: A Novice-Oriented Live Programming Environment with Always-On Run-Time Value VisualizationsabstractVisualizations of run-time program state help novices form proper mental models and debug their code. We push this technique to the extreme by posing the following question: What if a live programming environment for an imperative language always displays the entire history of all run-time values for all program variables all the time? To explore this question, we built a prototype live IDE called Omnicode ("Omniscient Code") that continually runs the user's Python code and uses a scatterplot matrix to visualize the entire history of all of its numerical values, along with meaningful numbers derived from other data types. To filter the visualizations and hone in on specific points of interest, the user can brush and link over the scatterplots or select portions of code. They can also zoom in to view detailed stack and heap visualizations at each execution step. An exploratory study on 10 novice programmers discovered that they found Omnicode to be useful for debugging, forming mental models, explaining their code to others, and discovering moments of serendipity that would not have been likely within an ordinary IDE. Hyeonsu B. Kang, Philip J. Guo |
UIST | 2 |
| 2017 | Torta: Generating Mixed-Media GUI and Command-Line App Tutorials Using Operating-System-Wide Activity TracingabstractTutorials are vital for helping people perform complex software-based tasks in domains such as programming, data science, system administration, and computational research. However, it is tedious to create detailed step-by-step tutorials for tasks that span multiple interrelated GUI and command-line applications. To address this challenge, we created Torta, an end-to-end system that automatically generates step-by-step GUI and command-line app tutorials by demonstration, provides an editor to trim, organize, and add validation criteria to these tutorials, and provides a web-based viewer that can validate step-level progress and automatically run certain steps. The core technical insight that underpins Torta is that combining operating-system-wide activity tracing and screencast recording makes it easier to generate mixed-media (text+video) tutorials that span multiple GUI and command-line apps. An exploratory study on 10 computer science teaching assistants (TAs) found that they all preferred the experience and results of using Torta to record programming and sysadmin tutorials relevant to classes they teach rather than manually writing tutorials. A follow-up study on 6 students found that they all preferred following the Torta tutorials created by those TAs over the manually-written versions. Alok Mysore, Philip J. Guo |
UIST | 2 |
| 2017 | DS.js: Turn Any Webpage into an Example-Centric Live Programming Environment for Learning Data ScienceabstractData science courses and tutorials have grown popular in recent years, yet they are still taught using production-grade programming tools (e.g., R, MATLAB, and Python IDEs) within desktop computing environments. Although powerful, these tools present high barriers to entry for novices, forcing them to grapple with the extrinsic complexities of software installation and configuration, data file management, data parsing, and Unix-like command-line interfaces. To lower the barrier for novices to get started with learning data science, we created DS.js, a bookmarklet that embeds a data science programming environment directly into any existing webpage. By transforming any webpage into an example-centric IDE, DS.js eliminates the aforementioned complexities of desktop-based environments and turns the entire web into a rich substrate for learning data science. DS.js automatically parses HTML tables and CSV/TSV data sets on the target webpage, attaches code editors to each data set, provides a data table manipulation and visualization API designed for novices, and gives instructional scaffolding in the form of bidirectional previews of how the user's code and data relate. Philip J. Guo |
UIST | 2 |
| 2017 | HappyFace: Identifying and predicting frustrating obstacles for learning programming at scaleabstractUnnecessary obstacles limit learning in cognitively-complex domains such as computer programming. With a lack of appropriate feedback mechanisms, novice programmers can experience frustration and disengage from the learning experience. In large-scale educational settings, the struggles of learners are often invisible to the learning infrastructure and learners have limited ability to seek help. In this paper, we perform a large-scale collection of code snippets from an online learn-to-code platform, Python Tutor, and collect a frustration rating through a light-weight learner feedback mechanism. We then devise a technique that can automatically identify sources of frustration based on participants labeling their frustration levels. We found 3 factors that best predicted novice programmers' frustration state: syntax errors, using niche language features, and understanding code with high complexity. Additionally, we found evidence that we could predict sources of frustration. Based on these results, we believe an embedded feedback mechanism can lead to future intervention systems. Ian Drosos, Philip J. Guo, Chris Parnin |
VL/HCC | 2 |
| 2016 | Understanding Conversational Programmers: A Perspective from the Software IndustryabstractRecent research suggests that some students learn to program with the goal of becoming conversational programmers: they want to develop programming literacy skills not to write code in the future but mainly to develop conversational skills and communicate better with developers and to improve their marketability. To investigate the existence of such a population of conversational programmers in practice, we surveyed professionals at a large multinational technology company who were not in software development roles. Based on 3151 survey responses from professionals who never or rarely wrote code, we found that a significant number of them (42.6%) had invested in learning programming on the job. While many of these respondents wanted to perform traditional end-user programming tasks (e.g., data analysis), we discovered that two top motivations for learning programming were to improve the efficacy of technical conversations and to acquire marketable skillsets. The main contribution of this work is in empirically establishing the existence and characteristics of conversational programmers in a large software development context. Parmit K. Chilana, Rishabh Singh, Philip J. Guo |
CHI | 3 |
| 2016 | Paradise unplugged: identifying barriers for female participation on stack overflowabstractIt is no secret that females engage less in programming fields than males. However, in online communities, such as Stack Overflow, this gender gap is even more extreme: only 5.8% of contributors are female. In this paper, we use a mixed-methods approach to identify contribution barriers females face in online communities. Through 22 semi-structured interviews with a spectrum of female users ranging from non-contributors to a top 100 ranked user of all time, we identified 14 barriers preventing them from contributing to Stack Overflow. We then conducted a survey with 1470 female and male developers to confirm which barriers are gender related or general problems for everyone. Females ranked five barriers significantly higher than males. A few of these include doubts in the level of expertise needed to contribute, feeling overwhelmed when competing with a large number of users, and limited awareness of site features. Still, there were other barriers that equally impacted all Stack Overflow users or affected particular groups, such as industry programmers. Finally, we describe several implications that may encourage increased participation in the Stack Overflow community across genders and other demographics. Denae Ford, Justin Smith 0001, Philip J. Guo, Chris Parnin |
SIGSOFT FSE | 3 |
| 2015 | Wait-Learning: Leveraging Wait Time for Second Language EducationabstractCompeting priorities in daily life make it difficult for those with a casual interest in learning to set aside time for regular practice. In this paper, we explore wait-learning: leveraging brief moments of waiting during a person's existing conversations for second language vocabulary practice, even if the conversation happens in the native language. We present an augmented version of instant messaging, WaitChatter, that supports the notion of wait-learning by displaying contextually relevant foreign language vocabulary and micro-quizzes just-in-time while the user awaits a response from her conversant. Through a two week field study of WaitChatter with 20 people, we found that users were able to learn 57 new words on average during casual instant messaging. Furthermore, we found that users were most receptive to learning opportunities immediately after sending a chat message, and that this timing may be critical given user tendency to multi-task during waiting periods. Carrie J. Cai, Philip J. Guo, James R. Glass, Rob Miller 0001 |
CHI | 2 |
| 2015 | How High School, College, and Online Students Differentially Engage with an Interactive Digital Textbook
Jeremy Warner, John Doorenbos, Bradley N. Miller, Philip J. Guo |
EDM | 4 |
| 2015 | Codeopticon: Real-Time, One-To-Many Human Tutoring for Computer ProgrammingabstractOne-on-one tutoring from a human expert is an effective way for novices to overcome learning barriers in complex domains such as computer programming. But there are usually far fewer experts than learners. To enable a single expert to help more learners at once, we built Codeopticon, an interface that enables a programming tutor to monitor and chat with dozens of learners in real time. Each learner codes in a workspace that consists of an editor, compiler, and visual debugger. The tutor sees a real-time view of each learner's actions on a dashboard, with each learner's workspace summarized in a tile. At a glance, the tutor can see how learners are editing and debugging their code, and what errors they are encountering. The dashboard automatically reshuffles tiles so that the most active learners are always in the tutor's main field of view. When the tutor sees that a particular learner needs help, they can open an embedded chat window to start a one-on-one conversation. A user study showed that 8 first-time Codeopticon users successfully tutored anonymous learners from 54 countries in a naturalistic online setting. On average, in a 30-minute session, each tutor monitored 226 learners, started 12 conversations, exchanged 47 chats, and helped 2.4 learners. Philip J. Guo |
UIST | 1 |
| 2015 | Perceptions of non-CS majors in intro programming: The rise of the conversational programmerabstractDespite the enthusiasm and initiatives for making programming accessible to students outside Computer Science (CS), unfortunately, there are still many unanswered questions about how we should be teaching programming to engineers, scientists, artists or other non-CS majors. We present an in-depth case study of first-year management engineering students enrolled in a required introductory programming course at a large North American university. Based on an inductive analysis of one-on-one interviews, surveys, and weekly observations, we provide insights into students' motivations, career goals, perceptions of programming, and reactions to the Java and Processing languages. One of our key findings is that between the traditional classification of non-programmers vs. programmers, there exists a category of conversational programmers who do not necessarily want to be professional programmers or even end-user programmers, but want to learn programming so that they can speak in the “programmer's language” and improve their perceived job marketability in the software industry. Parmit K. Chilana, Celena Alcock, Shruti Dembla, Anson Ho, Ada Hurst, Brett Armstrong, Philip J. Guo |
VL/HCC | 7 |
| 2015 | Codepourri: Creating visual coding tutorials using a volunteer crowd of learnersabstractA common way to learn is by studying written step-by-step tutorials such as worked examples. However, tutorials for computer programming can be tedious to create since a static text-based format cannot convey what happens as code executes. We created a system called Codepourri that enables people to easily create visual coding tutorials by annotating steps in an automatically-generated program visualization. Using Codepourri, we developed a novel crowdsourcing workflow where learners who are visiting an educational Web site (www. pythontutor.com) collectively create a tutorial by annotating execution steps in a piece of code and then voting on the best annotations. Since there are far more learners than experts, using learners as a crowd is a potentially more scalable way of creating tutorials. Our experiments with 4 expert judges and 101 learners adding 145 raw annotations to two pieces of textbook Python code show the learner crowd's annotations to be accurate, informative, and containing some insights that even experts missed. Mitchell L. Gordon, Philip J. Guo |
VL/HCC | 2 |
| 2015 | Codechella: Multi-user program visualizations for real-time tutoring and collaborative learningabstractAn effective way to learn computer programming is to sit side-by-side in front of the same computer with a tutor or peer, write code together, and then discuss what happens as the code executes. To bring this kind of in-person interaction to an online setting, we have developed Codechella, a multi-user Web-based program visualization system that enables multiple people to collaboratively write code together, explore an automatically-generated visualization of its execution state using multiple mouse cursors, and chat via an embedded text box. In nine months of live deployment on an educational website - www.pythontutor.com -people from 296 cities across 40 countries participated in 299 Codechella sessions for both tutoring and collaborative learning. 57% of sessions connected participants from different cities, and 12% from different countries. Participants actively engaged with the program visualizations while chatting, showed affective exchanges such as encouragement and banter, and indicated signs of learning at the lower three levels of Bloom's taxonomy: remembering, understanding, and applying knowledge. Philip J. Guo, Jeffery White, Renan Zanelatto |
VL/HCC | 1 |
| 2015 | Toward a domain-specific visual discussion forum for learning computer programming: An empirical study of a popular MOOC forumabstractOnline discussion forums are one of the most ubiquitous kinds of resources for people who are learning computer programming. However, their user interface - a hierarchy of textual threads - has not changed much in the past four decades. We argue that generic forum interfaces are cumbersome for learning programming and that there is a need for a domain-specific visual discussion forum for programming. We support this argument with an empirical study of all 5,377 forum threads in Introduction to Computer Science and Programming Using Python, a popular edX MOOC. Specifically, we investigated how forum participants were hampered by its text-based format. Most notably, people often wanted to discuss questions about dynamic execution state - what happens “under the hood” as the computer runs code. We propose that a better forum for learning programming should be visual and domain-specific, integrating automatically-generated visualizations of execution state and enabling inline annotations of source code and output. Joyce Zhu, Jeremy Warner, Mitchell L. Gordon, Jeffery White, Renan Zanelatto, Philip J. Guo |
VL/HCC | 6 |
| 2015 | OverCode: Visualizing Variation in Student Solutions to Programming Problems at ScaleabstractIn MOOCs, a single programming exercise may produce thousands of solutions from learners. Understanding solution variation is important for providing appropriate feedback to students at scale. The wide variation among these solutions can be a source of pedagogically valuable examples and can be used to refine the autograder for the exercise by exposing corner cases. We present OverCode, a system for visualizing and exploring thousands of programming solutions. OverCode uses both static and dynamic analysis to cluster similar solutions, and lets teachers further filter and cluster solutions based on different criteria. We evaluated OverCode against a nonclustering baseline in a within-subjects study with 24 teaching assistants and found that the OverCode interface allows teachers to more quickly develop a high-level view of students' understanding and misconceptions, and to provide feedback that is relevant to more students' solutions. Elena L. Glassman, Jeremy Scott, Rishabh Singh, Philip J. Guo, Rob Miller 0001 |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2014 | Crowdsourcing step-by-step information extraction to enhance existing how-to videosabstractMillions of learners today use how-to videos to master new skills in a variety of domains. But browsing such videos is often tedious and inefficient because video player interfaces are not optimized for the unique step-by-step structure of such videos. This research aims to improve the learning experience of existing how-to videos with step-by-step annotations. Juho Kim 0001, Phu Tran Nguyen, Sarah A. Weir, Philip J. Guo, Rob Miller 0001, Krzysztof Z. Gajos |
CHI | 4 |
| 2014 | How video production affects student engagement: an empirical study of MOOC videosabstractVideos are a widely-used kind of resource for online learning. This paper presents an empirical study of how video production decisions affect student engagement in online educational videos. To our knowledge, ours is the largest-scale study of video engagement to date, using data from 6.9 million video watching sessions across four courses on the edX MOOC platform. We measure engagement by how long students are watching each video, and whether they attempt to answer post-video assessment problems. Philip J. Guo, Juho Kim 0001, Rob Rubin |
L@S | 1 |
| 2014 | Demographic differences in how students navigate through MOOCsabstractThe current generation of Massive Open Online Courses (MOOCs) attract a diverse student audience from all age groups and over 196 countries around the world. Researchers, educators, and the general public have recently become interested in how the learning experience in MOOCs differs from that in traditional courses. A major component of the learning experience is how students navigate through course content. Philip J. Guo, Katharina Reinecke |
L@S | 1 |
| 2014 | Understanding in-video dropouts and interaction peaks inonline lecture videosabstractWith thousands of learners watching the same online lecture videos, analyzing video watching patterns provides a unique opportunity to understand how students learn with videos. This paper reports a large-scale analysis of in-video dropout and peaks in viewership and student activity, using second-by-second user interaction data from 862 videos in four Massive Open Online Courses (MOOCs) on edX. We find higher dropout rates in longer videos, re-watching sessions (vs first-time), and tutorials (vs lectures). Peaks in re-watching sessions and play events indicate points of interest and confusion. Results show that tutorials (vs lectures) and re-watching sessions (vs first-time) lead to more frequent and sharper peaks. In attempting to reason why peaks occur by sampling 80 videos, we observe that 61% of the peaks accompany visual transitions in the video, e.g., a slide view to a classroom view. Based on this observation, we identify five student activity patterns that can explain peaks: starting from the beginning of a new material, returning to missed content, following a tutorial step, replaying a brief segment, and repeating a non-visual explanation. Our analysis has design implications for video authoring, editing, and interface design, providing a richer understanding of video learning on MOOCs. Juho Kim 0001, Philip J. Guo, Daniel T. Seaton, Piotr Mitros, Krzysztof Z. Gajos, Rob Miller 0001 |
L@S | 2 |
| 2014 | Modeling programming knowledge for mentoring at scaleabstractIn large programming classes, MOOCs or online communities, it is challenging to find peers and mentors to help with learning specific programming concepts. In this paper we present first steps towards an automated, scalable system for matching learners with Python programmers who have expertise in different areas. The learner matching system builds a knowledge model for each programmer by analyzing their authored code and extracting features that capture domain knowledge and style. We demonstrate the feasibility of a simple model that counts the references to modules from the standard library and Python Package Index in a programmers' code. We also show that programmers exhibit self-selection using which we can extract the modules a programmer is best at, even though we may not have all of their code. In our future work we aim to extend the model to encapsulate more features, and apply it for skill matching in a programming class as well as personalizing answers on StackOverflow. Anvisha H. Pai, Philip J. Guo, Rob Miller 0001 |
L@S | 2 |
| 2014 | Data-driven interaction techniques for improving navigation of educational videosabstractWith an unprecedented scale of learners watching educational videos on online platforms such as MOOCs and YouTube, there is an opportunity to incorporate data generated from their interactions into the design of novel video interaction techniques. Interaction data has the potential to help not only instructors to improve their videos, but also to enrich the learning experience of educational video watchers. This paper explores the design space of data-driven interaction techniques for educational video navigation. We introduce a set of techniques that augment existing video interface widgets, including: a 2D video timeline with an embedded visualization of collective navigation traces; dynamic and non-linear timeline scrubbing; data-enhanced transcript search and keyword summary; automatic display of relevant still frames next to the video; and a visual summary representing points with high learner activity. To evaluate the feasibility of the techniques, we ran a laboratory user study with simulated learning tasks. Participants rated watching lecture videos with interaction data to be efficient and useful in completing the tasks. However, no significant differences were found in task performance, suggesting that interaction data may not always align with moment-by-moment information needs during the tasks. Juho Kim 0001, Philip J. Guo, Carrie J. Cai, Shang-Wen Li 0001, Krzysztof Z. Gajos, Rob Miller 0001 |
UIST | 2 |
| 2014 | A direct manipulation language for explaining algorithmsabstractInstructors typically explain algorithms in computer science by tracing their behavior, often on blackboards, sometimes with algorithm visualizations. Using blackboards can be tedious because they do not facilitate manipulation of the drawing, while visualizations often operate at the wrong level of abstraction or must be laboriously hand-coded for each algorithm. In response, we present a direct manipulation (DM) language for explaining algorithms by manipulating visualized data structures. The language maps DM gestures onto primitive program behaviors that occur in commonly taught algorithms. We performed an initial evaluation of the DM language on teaching assistants of an undergraduate algorithms class, who found the language easier to use and more helpful for explaining algorithms than a standard drawing application (GIMP). Jeremy Scott, Philip J. Guo, Randall Davis |
VL/HCC | 2 |
| 2013 | Online python tutor: embeddable web-based program visualization for cs educationabstractThis paper presents Online Python Tutor, a web-based program visualization tool for Python, which is becoming a popular language for teaching introductory CS courses. Using this tool, teachers and students can write Python programs directly in the web browser (without installing any plugins), step forwards and backwards through execution to view the run-time state of data structures, and share their program visualizations on the web. In the past three years, over 200,000 people have used Online Python Tutor to visualize their programs. In addition, instructors in a dozen universities such as UC Berkeley, MIT, the University of Washington, and the University of Waterloo have used it in their CS1 courses. Finally, Online Python Tutor visualizations have been embedded within three web-based digital Python textbook projects, which collectively attract around 16,000 viewers per month and are being used in at least 25 universities. Online Python Tutor is free and open source software, available at pythontutor.com. Philip J. Guo |
SIGCSE | 1 |
| 2012 | Characterizing and predicting which bugs get reopenedabstractFixing bugs is an important part of the software development process. An underlying aspect is the effectiveness of fixes: if a fair number of fixed bugs are reopened, it could indicate instability in the software system. To the best of our knowledge there has been on little prior work on understanding the dynamics of bug reopens. Towards that end, in this paper, we characterize when bug reports are reopened by using the Microsoft Windows operating system project as an empirical case study. Our analysis is based on a mixed-methods approach. First, we categorize the primary reasons for reopens based on a survey of 358 Microsoft employees. We then reinforce these results with a large-scale quantitative study of Windows bug reports, focusing on factors related to bug report edits and relationships between people involved in handling the bug. Finally, we build statistical models to describe the impact of various metrics on reopening bugs ranging from the reputation of the opener to how the bug was found. Thomas Zimmermann 0001, Nachiappan Nagappan, Philip J. Guo, Brendan Murphy |
ICSE | 3 |
| 2012 | HAMPI: A solver for word equations over strings, regular expressions, and context-free grammarsabstractMany automatic testing, analysis, and verification techniques for programs can be effectively reduced to a constraint-generation phase followed by a constraint-solving phase. This separation of concerns often leads to more effective and maintainable software reliability tools. The increasing efficiency of off-the-shelf constraint solvers makes this approach even more compelling. However, there are few effective and sufficiently expressive off-the-shelf solvers for string constraints generated by analysis of string-manipulating programs, so researchers end up implementing their own ad-hoc solvers. To fulfill this need, we designed and implemented Hampi, a solver for string constraints over bounded string variables. Users of Hampi specify constraints using regular expressions, context-free grammars, equality between string terms, and typical string operations such as concatenation and substring extraction. Hampi then finds a string that satisfies all the constraints or reports that the constraints are unsatisfiable. We demonstrate Hampi's expressiveness and efficiency by applying it to program analysis and automated testing. We used Hampi in static and dynamic analyses for finding SQL injection vulnerabilities in Web applications with hundreds of thousands of lines of code. We also used Hampi in the context of automated bug finding in C programs using dynamic systematic testing (also known as concolic testing). We then compared Hampi with another string solver, CFGAnalyzer, and show that Hampi is several times faster. Hampi's source code, documentation, and experimental data are available at http://people.csail.mit.edu/akiezun/hampi 1 Adam Kiezun, Vijay Ganesh 0001, Shay Artzi, Philip J. Guo, Pieter Hooimeijer, Michael D. Ernst |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2011 | HAMPI: A String Solver for Testing, Analysis and Vulnerability Detection
Vijay Ganesh 0001, Adam Kiezun, Shay Artzi, Philip J. Guo, Pieter Hooimeijer, Michael D. Ernst |
CAV | 4 |
| 2011 | "Not my bug!" and other reasons for software bug report reassignmentsabstractBug reporting/fixing is an important social part of the soft-ware development process. The bug-fixing process inher-ently has strong inter-personal dynamics at play, especially in how to find the optimal person to handle a bug report. Bug report reassignments, which are a common part of the bug-fixing process, have rarely been studied. Philip J. Guo, Thomas Zimmermann 0001, Nachiappan Nagappan, Brendan Murphy |
CSCW | 1 |
| 2011 | Using automatic persistent memoization to facilitate data analysis scriptingabstractProgrammers across a wide range of disciplines (e.g., bioinformatics, neuroscience, econometrics, finance, data mining, information retrieval, machine learning) write scripts to parse, transform, process, and extract insights from data. To speed up iteration times, they split their analyses into stages and write extra code to save the intermediate results of each stage to files so that those results do not have to be re-computed in every subsequent run. As they explore and refine hypotheses, their scripts often create and process lots of intermediate data files. They need to properly manage the myriad of dependencies between their code and data files, or else their analyses will produce incorrect results. Philip J. Guo, Dawson R. Engler |
ISSTA | 1 |
| 2011 | CDE: Run Any Linux Application On-Demand Without Installation
Philip J. Guo |
LISA | 1 |
| 2011 | Proactive wrangling: mixed-initiative end-user programming of data transformation scriptsabstractAnalysts regularly wrangle data into a form suitable for computational tools through a tedious process that delays more substantive analysis. While interactive tools can assist data transformation, analysts must still conceptualize the desired output state, formulate a transformation strategy, and specify complex transforms. We present a model to proactively suggest data transforms which map input data to a relational format expected by analysis tools. To guide search through the space of transforms, we propose a metric that scores tables according to type homogeneity, sparsity and the presence of delimiters. When compared to "ideal" hand-crafted transformations, our model suggests over half of the needed steps; in these cases the top-ranked suggestion is preferred 77% of the time. User study results indicate that suggestions produced by our model can assist analysts' transformation tasks, but that users do not always value proactive assistance, instead preferring to maintain the initiative. We discuss some implications of these results for mixed-initiative interfaces. Philip J. Guo, Sean Kandel, Joseph M. Hellerstein, Jeffrey Heer |
UIST | 1 |
| 2011 | CDE: Using System Call Interposition to Automatically Create Portable Software Packages
Philip J. Guo, Dawson R. Engler |
USENIX ATC | 1 |
| 2010 | Characterizing and predicting which bugs get fixed: an empirical study of Microsoft WindowsabstractWe performed an empirical study to characterize factors that affect which bugs get fixed in Windows Vista and Windows 7, focusing on factors related to bug report edits and relationships between people involved in handling the bug. We found that bugs reported by people with better reputations were more likely to get fixed, as were bugs handled by people on the same team and working in geographical proximity. We reinforce these quantitative results with survey feedback from 358 Microsoft employees who were involved in Windows bugs. Survey respondents also mentioned additional qualitative influences on bug fixing, such as the importance of seniority and interpersonal skills of the bug reporter. Philip J. Guo, Thomas Zimmermann 0001, Nachiappan Nagappan, Brendan Murphy |
ICSE (1) | 1 |
| 2009 | Two studies of opportunistic programming: interleaving web foraging, learning, and writing codeabstractThis paper investigates the role of online resources in problem solving. We look specifically at how programmers - an exemplar form of knowledge workers - opportunistically interleave Web foraging, learning, and writing code. We describe two studies of how programmers use online resources. The first, conducted in the lab, observed participants' Web use while building an online chat room. We found that programmers leverage online resources with a range of intentions: They engage in just-in-time learning of new skills and approaches, clarify and extend their existing knowledge, and remind themselves of details deemed not worth remembering. The results also suggest that queries for different purposes have different styles and durations. Do programmers' queries "in the wild" have the same range of intentions, or is this result an artifact of the particular lab setting? We analyzed a month of queries to an online programming portal, examining the lexical structure, refinements made, and result pages visited. Here we also saw traits that suggest the Web is being used for learning and reminding. These results contribute to a theory of online resource usage in programming, and suggest opportunities for tools to facilitate online knowledge work. Joel Brandt, Philip J. Guo, Joel Lewenstein, Mira Dontcheva, Scott R. Klemmer |
CHI | 2 |
| 2009 | Automatic creation of SQL Injection and cross-site scripting attacksabstractWe present a technique for finding security vulnerabilities in Web applications. SQL Injection (SQLI) and cross-site scripting (XSS) attacks are widespread forms of attack in which the attacker crafts the input to the application to access or modify user data and execute malicious code. In the most serious attacks (called second-order, or persistent, XSS), an attacker can corrupt a database so as to cause subsequent users to execute malicious code. Adam Kiezun, Philip J. Guo, Karthick Jayaraman, Michael D. Ernst |
ICSE | 2 |
| 2009 | HAMPI: a solver for string constraintsabstractMany automatic testing, analysis, and verification techniques for programs can be effectively reduced to a constraint generation phase followed by a constraint-solving phase. This separation of concerns often leads to more effective and maintainable tools. The increasing efficiency of off-the-shelf constraint solvers makes this approach even more compelling. However, there are few effective and sufficiently expressive off-the-shelf solvers for string constraints generated by analysis techniques for string-manipulating programs. Adam Kiezun, Vijay Ganesh 0001, Philip J. Guo, Pieter Hooimeijer, Michael D. Ernst |
ISSTA | 3 |
| 2009 | Linux Kernel Developer Responses to Static Analysis Bug Reports
Philip J. Guo, Dawson R. Engler |
USENIX ATC | 1 |
| 2007 | The Daikon system for dynamic detection of likely invariants
Michael D. Ernst, Jeff H. Perkins, Philip J. Guo, Stephen McCamant, Carlos Pacheco, Matthew S. Tschantz, Chen Xiao |
Sci. Comput. Program. | 3 |
| 2006 | Inference and enforcement of data structure consistency specificationsabstractCorrupt data structures are an important cause of unacceptable program execution. Data structure repair (which eliminates inconsistencies by updating corrupt data structures to conform to consistency constraints) promises to enable many programs to continue to execute acceptably in the face of otherwise fatal data structure corruption errors. A key issue is obtaining an accurate and comprehensive data structure consistency specification. We present a new technique for obtaining data structure consistency specifications for data structure repair. Instead of requiring the developer to manually generate such specifications, our approach automatically generates candidate data structure consistency properties using the Daikon invariant detection tool. The developer then reviews these properties, potentially rejecting or generalizing overly specific properties to obtain a specification suitable for automatic enforcement via data structure repair. We have implemented this approach and applied it to three sizable benchmark programs: CTAS (an air-traffic control system), BIND (a widely-used Internet name server) and Freeciv (an interactive game). Our results indicate that (1) automatic constraint generation produces constraints that enable programs to execute successfully through data structure consistency errors, (2) compared to manual specification, automatic generation can produce more comprehensive sets of constraints that cover a larger range of data structure consistency properties, and (3) reviewing the properties is relatively straightforward and requires substantially less programmer effort than manual generation, primarily because it reduces the need to examine the program text to understand its operation and extract the relevant consistency constraints. Moreover, when evaluated by a hostile third party "Red Team" contracted to evaluate the effectiveness of the technique, our data structure inference and enforcement tools successfully prevented several otherwise fatal attacks. Brian Demsky, Michael D. Ernst, Philip J. Guo, Stephen McCamant, Jeff H. Perkins, Martin C. Rinard |
ISSTA | 3 |
| 2006 | Dynamic inference of abstract typesabstractAn abstract type groups variables that are used for related purposes in a program. We describe a dynamic unification-based analysis for inferring abstract types. Initially, each run-time value gets a unique abstract type. A run-time interaction among values indicates that they have the same abstract type, so their abstract types are unified. Also at run time, abstract types for variables are accumulated from abstract types for values. The notion of interaction may be customized, permitting the analysis to compute finer or coarser abstract types; these different notions of abstract type are useful for different tasks. We have implemented the analysis for compiled x86 binaries and for Java bytecodes. Our experiments indicate that the inferred abstract types are useful for program comprehension, improve both the results and the run time of a follow-on program analysis, and are more precise than the output of a comparable static analysis, without suffering from overfitting. Philip J. Guo, Jeff H. Perkins, Stephen McCamant, Michael D. Ernst |
ISSTA | 1 |