EDBT 2026 Demo / reviewers in the wild / expert
Elena L. Glassman
dblp:118/6231 · also Elena Leah Glassman
· DBLP profile ↗
47ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0001-5178-3496ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 37 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Systems, architecture and hardware · 7 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Paradigm for Creative OwnershipabstractAs generative AI tools become embedded in creative practice, questions of ownership in co-creative contexts are pressing. Yet studies of human-AI collaboration often invoke "ownership" without definition: sometimes conflating it with other concepts, and other times leaving interpretation to participants. This inconsistency makes findings difficult to compare across or even within studies. We introduce a framework of creative ownership comprising three dimensions - Person, Process, and System - each with three subdimensions, offering a shared language for both system design and HCI research. In semi-structured interviews with 21 creative professionals, we found that participants’ initial references to ownership (e.g., embodiment, control, concept) were fully encompassed by the framework, demonstrating its coverage. Once introduced, however, they also articulated and prioritized the remaining subdimensions, underscoring how the framework expands reflection and enables richer insights. Our contributions include 1) the framework, 2) a web-based visualization tool, and 3) empirical findings on its utility. Tejaswi Polimetla, Katy Ilonka Gero, Elena L. Glassman |
CHI | 3 |
| 2026 | How Notations Evolve: A Historical Analysis with Implications for Supporting User-Defined AbstractionsabstractTraditional human-computer interaction takes place through formally-specified systems like structured UIs and programming languages. Recent AI systems promise a new set of informal interactions with computers through natural language and other notational forms. These informal interactions can then lead to formal representations, but depend upon pre-existing formalisms known to both humans and AI. What about novel formalisms and notations? How are new abstractions created, evolved, and incrementally formalized over time—and how might new systems, in turn, be explicitly designed to support these processes? We conduct a comparative historical analysis of notation development to identify some relevant characteristics. These include three social stages of notation development: invention & incubation, dispersion & divergence, and institutionalization & sanctification, as well as three functional stages: descriptive, generative, and evaluative. Within and across these stages, we detail several patterns, such as the role of linking and grounding metaphors, dimensions of meaningful variation, and analogical alignment. Finally, we offer some implications for design. Jingyue Zhang, J. D. Zamfirescu-Pereira, Elena L. Glassman, Damien Masson, Ian Arawjo |
CHI | 3 |
| 2025 | CorpusStudio: Surfacing Emergent Patterns In A Corpus Of Prior Work While WritingabstractMany communities, including the scientific community, develop implicit writing norms. Understanding them is crucial for effective communication with that community. Writers gradually develop an implicit understanding of norms by reading papers and receiving feedback on their writing. However, it is difficult to both externalize this knowledge and apply it to one's own writing. We propose two new writing support concepts that reify document and sentence-level patterns in a given text corpus: (1) an ordered distribution over section titles and (2) given the user's draft and cursor location, many retrieved contextually relevant sentences. Recurring words in the latter are algorithmically highlighted to help users see any emergent norms. Study results (N=16) show that participants revised the structure and content using these concepts, gaining confidence in aligning with or breaking norms after reviewing many examples. These results demonstrate the value of reifying distributions over other authors' writing choices during the writing process. Hai Dang, Chelse Swoopes, Daniel Buschek, Elena L. Glassman |
CHI | 4 |
| 2025 | Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive TheoriesabstractUser iterates on their label definitions 5 3 System generates counter examples that vary in single dimensions The system learns pattern rules 2 User labels generated counterexamples 4 Simret Araya Gebreegziabher, Yukun Yang 0008, Elena L. Glassman, Toby Jia-Jun Li |
CHI | 3 |
| 2025 | Creative Writers' Attitudes on Writing as Training Data for Large Language ModelsabstractPeer Reviewed Katy Ilonka Gero, Meera A. Desai, Carly Schnitzler, Nayun Eom, Jack Cushman, Elena L. Glassman |
CHI | 6 |
| 2025 | AbstractExplorer: Leveraging Structure-Mapping Theory to Enhance Comparative Close Reading at Scale
Ziwei Gu, Joyce Zhou, Ning-Er (Nina) Lei, Jonathan K. Kummerfeld, Mahmood Jasim, Narges Mahyar, Elena L. Glassman |
UIST | 7 |
| 2025 | Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale
Priyan Vaithilingam, Munyeong Kim, Frida-Cecilia Acosta-Parenteau, Amine Mhedhbi, Elena L. Glassman, Ian Arawjo |
UIST | 6 |
| 2024 | Imagining a Future of Designing with AI: Dynamic Grounding, Constructive Negotiation, and Sustainable MotivationabstractWe ideate a future design workflow that involves AI technology. Drawing from activity and communication theory, we attempt to isolate the new value that large AI models can provide design compared to past technologies. We arrive at three affordances—dynamic grounding, constructive negotiation, and sustainable motivation—that summarize latent qualities of natural language-enabled foundation models that, if explicitly designed for, can support the process of design. Through design fiction, we then imagine a future interface as a diegetic prototype, the story of Squirrel Game, that demonstrates each of our three affordances in a realistic usage scenario. Our design process, terminology, and diagrams aim to contribute to future discussions about the relative affordances of AI technology with regard to collaborating with human designers. Priyan Vaithilingam, Ian Arawjo, Elena L. Glassman |
Conference on Designing Interactive Systems | 3 |
| 2024 | ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingabstractEvaluating outputs of large language models (LLMs) is challenging, requiring making—and making sense of—many responses. Yet tools that go beyond basic prompting tend to require knowledge of programming APIs, focus on narrow domains, or are closed-source. We present ChainForge, an open-source visual toolkit for prompt engineering and on-demand hypothesis testing of text generation LLMs. ChainForge provides a graphical interface for comparison of responses across models and prompt variations. Our system was designed to support three tasks: model selection, prompt template design, and hypothesis testing (e.g., auditing). We released ChainForge early in its development and iterated on its design with academics and online users. Through in-lab and interview studies, we find that a range of people could use ChainForge to investigate hypotheses that matter to them, including in real-world settings. We identify three modes of prompt engineering and LLM hypothesis testing: opportunistic exploration, limited evaluation, and iterative refinement. Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, Elena L. Glassman |
CHI | 5 |
| 2024 | Supporting Sensemaking of Large Language Model Outputs at ScaleabstractLarge language models (LLMs) are capable of generating multiple responses to a single prompt, yet little effort has been expended to help end-users or system designers make use of this capability. In this paper, we explore how to present many LLM responses at once. We design five features, which include both pre-existing and novel methods for computing similarities and differences across textual documents, as well as how to render their outputs. We report on a controlled user study (n=24) and eight case studies evaluating these features and how they support users in different tasks. We find that the features support a wide variety of sensemaking tasks and even make tasks tractable that our participants previously considered to be too difficult to attempt. Finally, we present design guidelines to inform future explorations of new LLM interfaces. Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, Elena L. Glassman |
CHI | 5 |
| 2024 | An AI-Resilient Text Rendering Technique for Reading and Skimming DocumentsabstractReaders find text difficult to consume for many reasons. Summarization can address some of these difficulties, but introduce others, such as omitting, misrepresenting, or hallucinating information, which can be hard for a reader to notice. One approach to addressing this problem is to instead modify how the original text is rendered to make important information more salient. We introduce Grammar-Preserving Text Saliency Modulation (GP-TSM), a text rendering method with a novel means of identifying what to de-emphasize. Specifically, GP-TSM uses a recursive sentence compression method to identify successive levels of detail beyond the core meaning of a passage, which are de-emphasized by rendering words in successively lighter but still legible gray text. In a lab study (n=18), participants preferred GP-TSM over pre-existing word-level text rendering methods and were able to answer GRE reading comprehension questions more efficiently. Ziwei Gu, Ian Arawjo, Kenneth Li 0002, Jonathan K. Kummerfeld, Elena L. Glassman |
CHI | 5 |
| 2024 | DynaVis: Dynamically Synthesized UI Widgets for Visualization EditingabstractUsers often rely on GUIs to edit and interact with visualizations — a daunting task due to the large space of editing options. As a result, users are either overwhelmed by a complex UI or constrained by a custom UI with a tailored, fixed subset of options with limited editing flexibility. Natural Language Interfaces (NLIs) are emerging as a feasible alternative for users to specify edits. However, NLIs forgo the advantages of traditional GUI: the ability to explore and repeat edits and see instant visual feedback. Priyan Vaithilingam, Elena L. Glassman, Jeevana Priya Inala, Chenglong Wang 0005 |
CHI | 2 |
| 2024 | Amortizing Pragmatic Program Synthesis with RankingsabstractThe usage of Rational Speech Acts (RSA) framework has been successful in building pragmatic program synthesizers that return programs which, in addition to being logically consistent with user-generated examples, account for the fact that a user chooses their examples informatively. We present a general method of amortizing the slow, exact RSA synthesizer. Our method first compiles a communication dataset of partially ranked programs by querying the exact RSA synthesizer. It then distills a global ranking – a single, total ordering of all programs, to approximate the partial rankings from this dataset. This global ranking is then used at inference time to rank multiple logically consistent candidate programs generated from a fast, non-pragmatic synthesizer. Experiments on two program synthesis domains using our ranking method resulted in orders of magnitudes of speed ups compared to the exact RSA synthesizer, while being more accurate than a non-pragmatic synthesizer. Finally, we prove that in the special case of synthesis from a single example, this approximation is exact. Yewen Pu, Saujas Vaduguru, Priyan Vaithilingam, Elena L. Glassman, Daniel Fried |
ICML | 4 |
| 2024 | psudo: Exploring Multi-Channel Biomedical Image Data with Spatially and Perceptually Optimized PseudocoloringabstractOver the past century, multichannel fluorescence imaging has been pivotal in myriad scientific breakthroughs by enabling the spatial visualization of proteins within a biological sample. With the shift to digital methods and visualization software, experts can now flexibly pseudocolor and combine image channels, each corresponding to a different protein, to explore their spatial relationships. We thus propose psudo, an interactive system that allows users to create optimal color palettes for multichannel spatial data. In psudo, a novel optimization method generates palettes that maximize the perceptual differences between channels while mitigating confusing color blending in overlapping channels. We integrate this method into a system that allows users to explore multi-channel image data and compare and evaluate color palettes for their data. An interactive lensing approach provides on-demand feedback on channel overlap and a color confusion metric while giving context to the underlying channel values. Color palettes can be applied globally or, using the lens, to local regions of interest. We evaluate our palette optimization approach using three graphical perception tasks in a crowdsourced user study with 150 participants, showing that users are more accurate at discerning and comparing the underlying data using our approach. Additionally, we showcase psudo in a case study exploring the complex immune responses in cancer tissue data with a biologist. Simon Warchol, Jakob Troidl, Jeremy Muhlich, Robert Krüger, John Hoffer, Tica Lin, Johanna Beyer, Elena L. Glassman, Peter K. Sorger, Hanspeter Pfister |
Comput. Graph. Forum | 8 |
| 2024 | Reliability Criteria for News WebsitesabstractMisinformation poses a threat to democracy and to people’s health. Reliability criteria for news websites can help people identify misinformation. But despite their importance, there has been no empirically substantiated list of criteria for distinguishing reliable from unreliable news websites. We identify reliability criteria, describe how they are applied in practice, and compare them to prior work. Based on our analysis, we distinguish between manipulable and less manipulable criteria and compare politically diverse laypeople as end-users and journalists as expert users. We discuss 11 widely recognized criteria, including the following 6 criteria that are difficult to manipulate: content, political alignment, authors, professional standards, what sources are used, and a website’s reputation. Finally, we describe how technology may be able to support people in applying these criteria in practice to assess the reliability of websites. Hendrik Heuer, Elena L. Glassman |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2023 | PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule SynthesisabstractOver the years, the task of AI-assisted data annotation has seen remarkable advancements. However, a specific type of annotation task, the qualitative coding performed during thematic analysis, has characteristics that make effective human-AI collaboration difficult. Informed by a formative study, we designed PaTAT, a new AI-enabled tool that uses an interactive program synthesis approach to learn flexible and expressive patterns over user-annotated codes in real-time as users annotate data. To accommodate the ambiguous, uncertain, and iterative nature of thematic analysis, the use of user-interpretable patterns allows users to understand and validate what the system has learned, make direct fixes, and easily revise, split, or merge previously annotated codes. This new approach also helps human users to learn data characteristics and form new theories in addition to facilitating the “learning” of the AI model. PaTAT’s usefulness and effectiveness were evaluated in a lab user study. Simret Araya Gebreegziabher, Zheng Zhang 0043, Xiaohang Tang, Yihao Meng, Elena L. Glassman, Toby Jia-Jun Li |
CHI | 5 |
| 2023 | Where to Hide a Stolen Elephant: Leaps in Creative Writing with Multimodal Machine IntelligenceabstractWhile developing a story, novices and published writers alike have had to look outside themselves for inspiration. Language models have recently been able to generate text fluently, producing new stochastic narratives upon request. However, effectively integrating such capabilities with human cognitive faculties and creative processes remains challenging. We propose to investigate this integration with a multimodal writing support interface that offers writing suggestions textually, visually, and aurally. We conduct an extensive study that combines elicitation of prior expectations before writing, observation and semi-structured interviews during writing, and outcome evaluations after writing. Our results illustrate the individual and situational variation in machine-in-the-loop writing approaches, suggestion acceptance, and ways the system is helpful. Centrally, we report how participants perform integrative leaps , by which they do cognitive work to integrate suggestions of varying semantic relevance into their developing stories. We interpret these findings, offering modeling and design recommendations for future creative writing support technologies. Nikhil Singh 0003, Guillermo Bernal, Daria Savchenko, Elena L. Glassman |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2022 | A Comparative Evaluation of Interventions Against Misinformation: Augmenting the WHO ChecklistabstractDuring the COVID-19 pandemic, the World Health Organization provided a checklist to help people distinguish between accurate and misinformation. In controlled experiments in the United States and Germany, we investigated the utility of this ordered checklist and designed an interactive version to lower the cost of acting on checklist items. Across interventions, we observe non-trivial differences in participants’ performance in distinguishing accurate and misinformation between the two countries and discuss some possible reasons that may predict the future helpfulness of the checklist in different environments. The checklist item that provides source labels was most frequently followed and was considered most helpful. Based on our empirical findings, we recommend practitioners focus on providing source labels rather than interventions that support readers performing their own fact-checks, even though this recommendation may be influenced by the WHO’s chosen order. We discuss the complexity of providing such source labels and provide design recommendations. Hendrik Heuer, Elena L. Glassman |
CHI | 2 |
| 2022 | Revisiting Human-Robot Teaching and Learning Through the Lens of Human Concept LearningabstractWhen interacting with a robot, humans form con-ceptual models (of varying quality) which capture how the robot behaves. These conceptual models form just from watching or in-teracting with the robot, with or without conscious thought. Some methods select and present robot behaviors to improve human conceptual model formation; nonetheless, these methods and HRI more broadly have not yet consulted cognitive theories of human concept learning. These validated theories offer concrete design guidance to support humans in developing conceptual models more quickly, accurately, and flexibly. Specifically, Analogical Transfer Theory and the Variation Theory of Learning have been successfully deployed in other fields, and offer new insights for the HRI community about the selection and presentation of robot behaviors. Using these theories, we review and contextualize 35 prior works in human-robot teaching and learning, and we assess how these works incorporate or omit the design implications of these theories. From this review, we identify new opportunities for algorithms and interfaces to help humans more easily learn conceptual models of robot behaviors, which in turn can help humans become more effective robot teachers and collaborators. Serena Booth, Sanjana Sharma, Sarah Chung, Julie A. Shah, Elena L. Glassman |
HRI | 5 |
| 2022 | Do Explanations Increase the Effectiveness of AI-Crowd Generated Fake News Warnings?
Ziv Epstein, Nicolò Foppiani, Sophie Hilgard, Sanjana Sharma, Elena L. Glassman, David G. Rand |
ICWSM | 5 |
| 2022 | Concept-Annotated Examples for Library ComparisonabstractProgrammers often rely on online resources—such as code examples, documentation, blogs, and Q&A forums—to compare similar libraries and select the one most suitable for their own tasks and contexts. However, this comparison task is often done in an ad-hoc manner, which may result in suboptimal choices. Inspired by Analogical Learning and Variation Theory, we hypothesize that rendering many concept-annotated code examples from different libraries side-by-side can help programmers (1) develop a more comprehensive understanding of the libraries’ similarities and distinctions and (2) make more robust, appropriate library selections. We designed a novel interactive interface, ParaLib, and used it as a technical probe to explore to what extent many side-by-side concepted-annotated examples can facilitate the library comparison and selection process. A within-subjects user study with 20 programmers shows that, when using ParaLib, participants made more consistent, suitable library selections and provided more comprehensive summaries of libraries’ similarities and differences. Litao Yan, Miryung Kim, Björn Hartmann, Tianyi Zhang 0001, Elena L. Glassman |
UIST | 5 |
| 2021 | Interpretable Program SynthesisabstractProgram synthesis, which generates programs based on user-provided specifications, can be obscure and brittle: users have few ways to understand and recover from synthesis failures. We propose interpretable program synthesis, a novel approach that unveils the synthesis process and enables users to monitor and guide a synthesizer. We designed three representations that explain the underlying synthesis process with different levels of fidelity. We implemented an interpretable synthesizer for regular expressions and conducted a within-subjects study with eighteen participants on three challenging regex tasks. With interpretable synthesis, participants were able to reason about synthesis failures and provide strategic feedback, achieving a significantly higher success rate compared with a state-of-the-art synthesizer. In particular, participants with a high engagement tendency (as measured by NCS-6) preferred a deductive representation that shows the synthesis process in a search tree, while participants with a relatively low engagement tendency preferred an inductive representation that renders representative samples of programs enumerated during synthesis. Tianyi Zhang 0001, Zhiyang Chen 0004, Yuanli Zhu, Priyan Vaithilingam, Xinyu Wang 0006, Elena L. Glassman |
CHI | 6 |
| 2021 | Evaluating the Interpretability of Generative Models by Interactive Reconstruction
Andrew Slavin Ross, Nina Chen, Elisa Zhao Hang, Elena L. Glassman, Finale Doshi-Velez |
CHI | 4 |
| 2021 | Visualizing Examples of Deep Neural Networks at ScaleabstractMany programmers want to use deep learning due to its superior accuracy in many challenging domains. Yet our formative study with ten programmers indicated that, when constructing their own deep neural networks (DNNs), they often had a difficult time choosing appropriate model structures and hyperparameter values. This paper presents ExampleNet—a novel interactive visualization system for exploring common and uncommon design choices in a large collection of open-source DNN projects. ExampleNet provides a holistic view of the distribution over model structures and hyperparameter settings in the corpus of DNNs, so users can easily filter the corpus down to projects tackling similar tasks and compare design choices made by others. We evaluated ExampleNet in a within-subjects study with sixteen participants. Compared with the control condition (i.e., online search), participants using ExampleNet were able to inspect more online examples, make more data-driven design decisions, and make fewer design mistakes. Litao Yan, Elena L. Glassman, Tianyi Zhang 0001 |
CHI | 2 |
| 2021 | Assuage: Assembly Synthesis Using A Guided ExplorationabstractAssembly programming is challenging, even for experts. Program synthesis, as an alternative to manual implementation, has the potential to enable both expert and non-expert users to generate programs in an automated fashion. However, current tools and techniques are unable to synthesize assembly programs larger than a few instructions. We present Assuage : ASsembly Synthesis Using A Guided Exploration, which is a parallel interactive assembly synthesizer that engages the user as an active collaborator, enabling synthesis to scale beyond current limits. Using Assuage, users can provide two types of semantically meaningful hints that expedite synthesis and allow for exploration of multiple possibilities simultaneously. Assuage exposes information about the underlying synthesis process using multiple representations to help users guide synthesis. We conducted a within-subjects study with twenty-one participants working on assembly programming tasks. With Assuage, participants with a wide range of expertise were able to achieve significantly higher success rates, perceived less subjective workload, and preferred the usefulness and usability of Assuage over a state of the art synthesis tool. Jingmei Hu, Priyan Vaithilingam, Stephen Chong, Margo I. Seltzer, Elena L. Glassman |
UIST | 5 |
| 2020 | Enabling Data-Driven API Design with Community Usage Data: A Need-Finding StudyabstractAPIs are becoming the fundamental building block of modern software and their usability is crucial to programming efficiency and software quality. Yet API designers find it hard to gather and interpret user feedback on their APIs. To close the gap, we interviewed 23 API designers from 6 companies and 11 open-source projects to understand their practices and needs. The primary way of gathering user feedback is through bug reports and peer reviews, as formal usability testing is prohibitively expensive to conduct in practice. Participants expressed a strong desire to gather real-world use cases and understand users' mental models, but there was a lack of tool support for such needs. In particular, participants were curious about where users got stuck, their workarounds, common mistakes, and unanticipated corner cases. We highlight several opportunities to address those unmet needs, including developing new mechanisms that systematically elicit users' mental models, building mining frameworks that identify recurring patterns beyond shallow statistics about API usage, and exploring alternative design choices made in similar libraries. Tianyi Zhang 0001, Björn Hartmann, Miryung Kim, Elena L. Glassman |
CHI | 4 |
| 2020 | Proxy tasks and subjective measures can be misleading in evaluating explainable AI systemsabstractExplainable artificially intelligent (XAI) systems form part of sociotechnical systems, e.g., human+AI teams tasked with making decisions. Yet, current XAI systems are rarely evaluated by measuring the performance of human+AI teams on actual decision-making tasks. We conducted two online experiments and one in-person think-aloud study to evaluate two currently common techniques for evaluating XAI systems: (1) using proxy, artificial tasks such as how well humans predict the AI's decision from the given explanations, and (2) using subjective measures of trust and preference as predictors of actual performance. The results of our experiments demonstrate that evaluations with proxy tasks did not predict the results of the evaluations with the actual decision-making tasks. Further, the subjective measures on evaluations with actual decision-making tasks did not predict the objective performance on those same tasks. Our results suggest that by employing misleading evaluation methods, our field may be inadvertently slowing its progress toward developing human+AI teams that can reliably perform better than humans or AIs alone. Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, Elena L. Glassman |
IUI | 4 |
| 2020 | Exempla gratis (E.G.): code examples for freeabstractModern software engineering often involves using many existing APIs, both open source and – in industrial coding environments– proprietary. Programmers reference documentation and code search tools to remind themselves of proper common usage patterns of APIs. However, high-quality API usage examples are computationally expensive to curate and maintain, and API usage examples retrieved from company-wide code search can be tedious to review. We present a tool, EG, that mines codebases and shows the common, idiomatic us-age examples for API methods. EG was integrated into Facebook’s internal code search tool for the Hack language and evaluated on open-source GitHub projects written in Python. EG was also compared against code search results and hand-written examples from a popular programming website called ProgramCreek. Compared with these two baselines, examples generated by EG are more succinct and representative with less extraneous statements. In addition, a survey with Facebook developers shows that EG examples are preferred in 97% of cases. Celeste Barnaby, Koushik Sen, Tianyi Zhang 0001, Elena L. Glassman, Satish Chandra 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2020 | Interactive Program Synthesis by Augmented ExamplesabstractProgramming-by-example (PBE) has become an increasingly popular component in software development tools, human-robot interaction, and end-user programming. A long-standing challenge in PBE is the inherent ambiguity in user-provided examples. This paper presents an interaction model to disambiguate user intent and reduce the cognitive load of understanding and validating synthesized programs. Our model provides two types of augmentations to user-given examples: 1) semantic augmentation where a user can specify how different aspects of an example should be treated by a synthesizer via light-weight annotations, and 2) data augmentation where the synthesizer generates additional examples to help the user understand and validate synthesized programs. We implement and demonstrate this interaction model in the domain of regular expressions, which is a popular mechanism for text processing and data wrangling and is often considered hard to master even for experienced programmers. A within-subjects user study with twelve participants shows that, compared with only inspecting and annotating synthesized programs, interacting with augmented examples significantly increases the success rate of finishing a programming task with less time and increases users? confidence of synthesized programs. Tianyi Zhang 0001, London Lowmanstone, Xinyu Wang 0006, Elena L. Glassman |
UIST | 4 |
| 2019 | Characterizing Developer Use of Automatically Generated PatchesabstractWe present a study that characterizes the way developers use automatically generated patches when fixing software defects. Our study tasked two groups of developers with repairing defects in C programs. Both groups were provided with the defective line of code. One was also provided with five automatically generated and validated patches, all of which modified the defective line of code, and one of which was correct. Contrary to our initial expectations, the group with access to the generated patches did not produce more correct patches and did not produce patches in less time. We characterize the main behaviors observed in experimental subjects: a focus on understanding the defect and the relationship of the patches to the original source code. Based on this characterization, we highlight various potentially productive directions for future developer-centric automatic patch generation systems. José Cambronero, Jiasi Shen 0001, Jürgen Cito, Elena L. Glassman, Martin C. Rinard |
VL/HCC | 4 |
| 2018 | Visualizing API Usage Examples at ScaleabstractUsing existing APIs properly is a key challenge in programming, given that libraries and APIs are increasing in number and complexity. Programmers often search for online code examples in Q&A forums and read tutorials and blog posts to learn how to use a given API. However, there are often a massive number of related code examples and it is difficult for a user to understand the commonalities and variances among them, while being able to drill down to concrete details. We introduce an interactive visualization for exploring a large collection of code examples mined from open-source repositories at scale. This visualization summarizes hundreds of code examples in one synthetic code skeleton with statistical distributions for canonicalized statements and structures enclosing an API call. We implemented this interactive visualization for a set of Java APIs and found that, in a lab study, it helped users (1) answer significantly more API usage questions correctly and comprehensively and (2) explore how other programmers have used an unfamiliar API. Elena L. Glassman, Tianyi Zhang 0001, Björn Hartmann, Miryung Kim |
CHI | 1 |
| 2018 | Interactive Extraction of Examples from Existing CodeabstractProgrammers frequently learn from examples produced and shared by other programmers. However, it can be challenging and time-consuming to produce concise, working code examples. We conducted a formative study where 12 participants made examples based on their own code. This revealed a key hurdle: making meaningful simplifications without introducing errors. Based on this insight, we designed a mixed-initiative tool, CodeScoop, to help programmers extract executable, simplified code from existing code. CodeScoop enables programmers to "scoop" out a relevant subset of code. Techniques include selectively including control structures and recording an execution trace that allows authors to substitute literal values for code and variables. In a controlled study with 19 participants, CodeScoop helped programmers extract executable code examples with the intended behavior more easily than with a standard code editor. Andrew Head, Elena L. Glassman, Björn Hartmann, Marti A. Hearst |
CHI | 2 |
| 2017 | Writing Reusable Code Feedback at Scale with Mixed-Initiative Program SynthesisabstractIn large introductory programming classes, teacher feedback on individual incorrect student submissions is often infeasible. Program synthesis techniques are capable of fixing student bugs and generating hints automatically, but they lack the deep domain knowledge of a teacher and can generate functionally correct but stylistically poor fixes. We introduce a mixed-initiative approach which combines teacher expertise with data-driven program synthesis techniques. We demonstrate our novel approach in two systems that use different interaction mechanisms. Our systems use program synthesis to learn bug-fixing code transformations and then cluster incorrect submissions by the transformations that correct them. The MistakeBrowser system learns transformations from examples of students fixing bugs in their own submissions. The FixPropagator system learns transformations from teachers fixing bugs in incorrect student submissions. Teachers can write feedback about a single submission or a cluster of submissions and propagate the feedback to all other submissions that can be fixed by the same transformation. Two studies suggest this approach helps teachers better understand student bugs and write reusable feedback that scales to a massive introductory programming classroom. Andrew Head, Elena L. Glassman, Gustavo Soares, Ryo Suzuki 0001, Lucas Figueredo, Loris D'Antoni, Björn Hartmann |
L@S | 2 |
| 2017 | Teamscope: Scalable Team Evaluation via Automated Metric Mining for Communication, Organization, Execution, and EvolutionabstractTeaching software development teams can be difficult to scale. Based on various cloud-based software development tools, Teamscope provides automated or semi-automated metrics to improve the scalability of a course with team projects. Metrics developed in Teamscope provide a synthesized view of a student team. Our preliminary results have shown the validity of these metrics. We also present a case study of applying metrics to teaching software development course in this paper. An Ju, Elena L. Glassman, Armando Fox |
L@S | 2 |
| 2017 | TraceDiff: Debugging unexpected code behavior using trace divergencesabstractRecent advances in program synthesis offer means to automatically debug student submissions and generate personalized feedback in massive programming classrooms. When automatically generating feedback for programming assignments, a key challenge is designing pedagogically useful hints that are as effective as the manual feedback given by teachers. Through an analysis of teachers' hint-giving practices in 132 online Q&A posts, we establish three design guidelines that an effective feedback design should follow. Based on these guidelines, we develop a feedback system that leverages both program synthesis and visualization techniques. Our system compares the dynamic code execution of both incorrect and fixed code and highlights how the error leads to a difference in behavior and where the incorrect code trace diverges from the expected solution. Results from our study suggest that our system enables students to detect and fix bugs that are not caught by students using another existing visual debugging tool. Ryo Suzuki 0001, Gustavo Soares, Andrew Head, Elena L. Glassman, Ruan Reis, Melina Mongiovi, Loris D'Antoni, Björn Hartmann |
VL/HCC | 4 |
| 2016 | Learnersourcing Personalized HintsabstractPersonalized support for students is a gold standard in education, but it scales poorly with the number of students. Prior work on learnersourcing presented an approach for learners to engage in human computation tasks while trying to learn a new skill. Our key insight is that students, through their own experience struggling with a particular problem, can become experts on the particular optimizations they implement or bugs they resolve. These students can then generate hints for fellow students based on their new expertise. We present workflows that harvest and organize students' collective knowledge and advice for helping fellow novices through design problems in engineering. Systems embodying each workflow were evaluated in the context of a college-level computer architecture class with an enrollment of more than two hundred students each semester. We show that, given our design choices, students can create helpful hints for their peers that augment or even replace teachers' personalized assistance, when that assistance is not available. Elena L. Glassman, Aaron Lin, Carrie J. Cai, Rob Miller 0001 |
CSCW | 1 |
| 2015 | Mudslide: A Spatially Anchored Census of Student Confusion for Online Lecture VideosabstractEducators have developed an effective technique to get feedback after in-person lectures, called "muddy cards." Students are given time to reflect and write the "muddiest" (least clear) point on an index card, to hand in as they leave class. This practice of assigning end-of-lecture reflection tasks to generate explicit student feedback is well suited for adaptation to the challenge of supporting feedback in online video lectures. We describe the design and evaluation of Mudslide, a prototype system that translates the practice of muddy cards into the realm of online lecture videos. Based on an in-lab study of students and teachers, we find that spatially contextualizing students' muddy point feedback with respect to particular lecture slides is advantageous to both students and teachers. We also reflect on further opportunities for enhancing this feedback method based on teachers' and students' experiences with our prototype. Elena L. Glassman, Juho Kim 0001, Andrés Monroy-Hernández, Meredith Ringel Morris |
CHI | 1 |
| 2015 | RIMES: Embedding Interactive Multimedia Exercises in Lecture VideosabstractTeachers in conventional classrooms often ask learners to express themselves and show their thought processes by speaking out loud, drawing on a whiteboard, or even using physical objects. Despite the pedagogical value of such activities, interactive exercises available in most online learning platforms are constrained to multiple-choice and short answer questions. We introduce RIMES, a system for easily authoring, recording, and reviewing interactive multimedia exercises embedded in lecture videos. With RIMES, teachers can prompt learners to record their responses to an activity using video, audio, and inking while watching lecture videos. Teachers can then review and interact with all the learners' responses in an aggregated gallery. We evaluated RIMES with 19 teachers and 25 students. Teachers created a diverse set of activities across multiple subjects that tested deep conceptual and procedural knowledge. Teachers found the exercises useful for capturing students' thought processes, identifying misconceptions, and engaging students with content. Juho Kim 0001, Elena L. Glassman, Andrés Monroy-Hernández, Meredith Ringel Morris |
CHI | 2 |
| 2015 | Learner-Sourcing in an Engineering Class at ScaleabstractTeaching computer architecture as a hands-on engineering course to approximately 250 MIT students per semester requires a large, dedicated teaching staff. This Spring, a shortened version of the course will be deployed on edX to a potentially far larger cohort of students, without additional teaching staff. To better support students, we have deployed developmental versions of three learner-sourcing systems to as many as 500 students. These systems harvest and organize students' collective knowledge about debugging and optimizing solutions. We plan to deploy and study the next iteration of these systems on edX this Spring. Elena L. Glassman, Christopher J. Terman, Rob Miller 0001 |
L@S | 1 |
| 2015 | Using and Designing Platforms for In Vivo Educational ExperimentsabstractIn contrast to typical laboratory experiments, the everyday use of online educational resources by large populations and the prevalence of software infrastructure for A/B testing leads us to consider how platforms can embed in vivo experiments that do not merely support research, but ensure practical improvements to their educational components. Examples are presented of randomized experimental comparisons conducted by subsets of the authors in three widely used online educational platforms -- Khan Academy, edX, and ASSISTments. We suggest design principles for platform technology to support randomized experiments that lead to practical improvements -- enabling Iterative Improvement and Collaborative Work -- and explain the benefit of their implementation by WPI co-authors in the ASSISTments platform. Joseph Jay Williams, Korinn S. Ostrow, Xiaolu Xiong, Elena L. Glassman, Juho Kim 0001, Samuel G. Maldonado, Na Li 0002, Justin Reich, Neil T. Heffernan |
L@S | 4 |
| 2015 | Foobaz: Variable Name Feedback for Student Code at ScaleabstractCurrent traditional feedback methods, such as hand-grading student code for substance and style, are labor intensive and do not scale. We created a user interface that addresses feedback at scale for a particular and important aspect of code quality: variable names. We built this user interface on top of an existing back-end that distinguishes variables by their behavior in the program. Therefore our interface not only allows teachers to comment on poor variable names, they can comment on names that mislead the reader about the variable's role in the program. We ran two user studies in which 10 teachers and 6 students created and received feedback, respectively. The interface helped teachers give personalized variable name feedback on thousands of student solutions from an edX introductory programming MOOC. In the second study, students composed solutions to the same programming assignments and immediately received personalized quizzes composed by teachers in the previous user study. Elena L. Glassman, Lyla Fischer, Jeremy Scott, Rob Miller 0001 |
UIST | 1 |
| 2015 | OverCode: Visualizing Variation in Student Solutions to Programming Problems at ScaleabstractIn MOOCs, a single programming exercise may produce thousands of solutions from learners. Understanding solution variation is important for providing appropriate feedback to students at scale. The wide variation among these solutions can be a source of pedagogically valuable examples and can be used to refine the autograder for the exercise by exposing corner cases. We present OverCode, a system for visualizing and exploring thousands of programming solutions. OverCode uses both static and dynamic analysis to cluster similar solutions, and lets teachers further filter and cluster solutions based on different criteria. We evaluated OverCode against a nonclustering baseline in a within-subjects study with 24 teaching assistants and found that the OverCode interface allows teachers to more quickly develop a high-level view of students' understanding and misconceptions, and to provide feedback that is relevant to more students' solutions. Elena L. Glassman, Jeremy Scott, Rishabh Singh, Philip J. Guo, Rob Miller 0001 |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2014 | Feature engineering for clustering student solutionsabstractOpen-ended homework problems such as coding assignments give students a broad range of freedom for the design of solutions. We aim to use the diversity in correct solutions to enhance student learning by automatically suggesting alternate solutions. Our approach is to perform a two-level hierarchical clustering of student solutions to first partition them based on the choice of algorithm and then partition solutions implementing the same algorithm based on low-level implementation details. Our initial investigations in domains of introductory programming and computer architecture demonstrate that we need two different classes of features to perform effective clustering at the two levels, namely abstract features and concrete features. Elena L. Glassman, Rishabh Singh, Rob Miller 0001 |
L@S | 1 |
| 2013 | Visualizing and classifying multiple solutions to engineering design problemsabstractIn engineering design courses, many problems have a specification that the student's implementation must meet, but give the student a broad range of freedom for the internal design of that implementation. There may be several distinct, correct strategies for solving them, some of which may be unknown to the teaching staff or intelligent tutor designer. Visualizing and classifying the multiple solutions that students generate in response to assigned engineering design problems will improve hints and answers to students' questions, whether they are provided by peers, staff, or automation. I log incremental snapshots of students' solutions as they progress toward correct and incorrect solutions. Initial investigations demonstrate that the choice of features to represent solutions is critical, and may be domain- or problem-dependent. Elena L. Glassman |
ICER | 1 |
| 2013 | Toward facilitating assistance to students attempting engineering design problemsabstractIn engineering design courses, many problems have a specification that the student's implementation must meet, but give the student a large range of freedom for the internal design of that implementation. There may be several distinct, correct strategies for solving them, some of which may be unknown to the teaching staff or intelligent tutor designer. When a student is pursuing an unrecognized strategy and begins to struggle, staff may redirect them, costing unnecessary work, and automated hint generators may offer unhelpful feedback. We have taken a first step toward discovering these alternate correct strategies by visualizing many student solutions together, using dynamic and static features of these solutions, so that the teaching staff can understand the space of correct strategies. This approach has been applied to two domains: an online Matlab programming challenge and an undergraduate computer architecture course. We discuss these initial investigations and pose discussion questions to the community about potential enhancement and application of this analysis. Elena L. Glassman, Ned Gulley, Rob Miller 0001 |
ICER | 1 |
| 2012 | Region of attraction estimation for a perching aircraft: A Lyapunov method exploiting barrier certificatesabstractDynamic perching maneuvers for fixed-wing aircraft are becoming increasingly plausible due to recent progress in perching using `micro-spines' mounted on tuned suspensions and, separately, on feedback motion planning techniques for post-stall maneuvering. In this paper, we bring these complementary techniques together by efficiently estimating the mechanical stability of the plane when it makes contact with a vertical surface; the resulting landing funnel can then be used in a feedback motion planning algorithm for the flight controller. We consider a simplified model of the perching dynamics and report an extension of the region of attraction techniques, using sums-of-squares optimization, which combines polynomial approximations of barrier constraints with the traditional Lyapunov methods to achieve tight estimation of the true region of attraction for the model. We demonstrate the new method on a variety of design parameters for the perching system, suggesting a potential use as a mechanical system or controller design tool. Elena L. Glassman, Alexis Lussier Desbiens, Mark M. Tobenkin, Mark R. Cutkosky, Russ Tedrake |
ICRA | 1 |
| 2010 | A quadratic regulator-based heuristic for rapidly exploring state spaceabstractKinodynamic planning algorithms like Rapidly-Exploring Randomized Trees (RRTs) hold the promise of finding feasible trajectories for rich dynamical systems with complex, nonconvex constraints. In practice, these algorithms perform very well on configuration space planning, but struggle to grow efficiently in systems with dynamics or differential constraints. This is due in part to the fact that the conventional distance metric, Euclidean distance, does not take into account system dynamics and constraints when identifying which node in the existing tree is capable of producing children closest to a given point in state space. We show that an affine quadratic regulator (AQR) design can be used to approximate the exact minimum-time distance pseudometric at a reasonable computational cost. We demonstrate improved exploration of the state spaces of the double integrator and simple pendulum when using this pseudometric within the RRT framework, but this improvement drops off as systems' nonlinearity and complexity increase. Future work includes exploring methods for approximating the exact minimum-time distance pseudometric that can reason about dynamics with higher-order terms. Elena L. Glassman, Russ Tedrake |
ICRA | 1 |