EDBT 2026 Demo / reviewers in the wild / expert
J. D. Zamfirescu-Pereira
dblp:239/8036
· DBLP profile ↗
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-5310-6728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 16 · 4 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Notations Evolve: A Historical Analysis with Implications for Supporting User-Defined AbstractionsabstractTraditional human-computer interaction takes place through formally-specified systems like structured UIs and programming languages. Recent AI systems promise a new set of informal interactions with computers through natural language and other notational forms. These informal interactions can then lead to formal representations, but depend upon pre-existing formalisms known to both humans and AI. What about novel formalisms and notations? How are new abstractions created, evolved, and incrementally formalized over time—and how might new systems, in turn, be explicitly designed to support these processes? We conduct a comparative historical analysis of notation development to identify some relevant characteristics. These include three social stages of notation development: invention & incubation, dispersion & divergence, and institutionalization & sanctification, as well as three functional stages: descriptive, generative, and evaluative. Within and across these stages, we detail several patterns, such as the role of linking and grounding metaphors, dimensions of meaningful variation, and analogical alignment. Finally, we offer some implications for design. Jingyue Zhang, J. D. Zamfirescu-Pereira, Elena L. Glassman, Damien Masson, Ian Arawjo |
CHI | 2 |
| 2026 | "Un-default" Behavior Tuning: Specifying Model Behavior outside the Norm with LLM Self-Playing and Self-ImprovingabstractSpecifying model behavior is challenging—especially when the desired behavior is unpopular relative to the model’s training data. Reversing the influence of massive training corpora is both time-consuming and costly, and such interventions are typically inaccessible to end users. While Large Language Models (LLMs) make it easier to write instructions using natural language, specifying unpopular behaviors remains a difficult task. Soya Park, J. D. Zamfirescu-Pereira, Chinmay Kulkarni 0001 |
IUI | 2 |
| 2026 | Instructors' Perspectives on LLM-Generated Programming Formative FeedbackabstractWe study instructor perspectives on LLM-generated programming feedback in an introductory Python course. LLM tutors predominantly offered debugging help, while human instructors preferred more diverse feedback types, including conceptual reminders, revisiting the problem, and examples. Cases where LLM tutor feedback diverged from human instructors' intent required major edits with different feedback types, while cases with closer alignment needed only minor changes with similar feedback types. Findings highlight the need for LLM tutors to reflect on instructor intent to ensure pedagogically aligned feedback. Rose Niousha, Samantha Boatright Smith, Abigail O'Neill, J. D. Zamfirescu-Pereira, John DeNero, Narges Norouzi |
SIGCSE (2) | 4 |
| 2026 | Misconception-Aware LLM Programming Tutor: Lessons Learned from Student-Tutor InteractionsabstractLarge Language Models (LLMs) are increasingly used as programming tutors, but their feedback is often generic and prone to solution leakage. To address these issues, we present MisconceptionTutor, which grounds feedback in common student misconceptions. Through both pre-deployment analyses and a real-classroom deployment, we find that even simple prompting frameworks can meaningfully steer tutor behavior to be more pedagogically oriented and noticeably more satisfying to students. Rose Niousha, Samantha Boatright Smith, Abigail O'Neill, J. D. Zamfirescu-Pereira, John DeNero, Narges Norouzi |
SIGCSE (2) | 4 |
| 2025 | Beyond Code Generation: LLM-supported Exploration of the Program Design SpaceabstractIn this work, we explore explicit Large Language Model (LLM)powered support for the iterative design of computer programs.Program design, like other design activity, is characterized by navigating a space of alternative problem formulations and associated solutions in an iterative fashion.LLMs are potentially powerful tools in helping this exploration; however, by default, code-generation LLMs deliver code that represents a particular point solution.This obscures the larger space of possible alternatives, many of which might be preferable to the LLM's default interpretation and its generated code.We contribute an IDE that supports program design through generating and showing new ways to frame problems alongside alternative solutions, tracking design decisions, and identifying implicit decisions made by either the programmer or the LLM.In a user study, we find that with our IDE, users combine and parallelize design phases to explore a broader design space-but also struggle to keep up with LLM-originated changes to code and other information overload.These findings suggest a core challenge for future IDEs that support program design through higher-level instructions given to LLM-based agents: carefully managing attention and deciding what information agents should surface to program designers and when. J. D. Zamfirescu-Pereira, Eunice Jun, Michael Terry, Qian Yang 0004, Björn Hartmann |
CHI | 1 |
| 2025 | From Code to Concepts: Textbook-Driven Knowledge Tracing with LLMs in CS1abstractGauging a student's understanding of course concepts, at an arbitrary point during a course, can be challenging. Standardized exams offer only a snapshot of performance rather than a deep understanding of progress. However, with Large Language Models (LLMs) now deployed at scale in CS1 courses, we can track multiple attempts from each student for every homework problem. This data provides insights into how students learn and deploy concepts over time, presenting a unique opportunity to rethink how we track changes in individual student knowledge. Traditional Knowledge Tracing (KT) methods often lack explainability and are computationally expensive. In contrast, our framework leverages an LLM to identify student progress on labeled, problem-level concepts from a student homework code submission. Our initial results show that the student's knowledge state can be dynamically updated. This knowledge state can then be used to provide more targeted, effective feedback and create tailored study materials. Abigail O'Neill, Samantha Boatright Smith, Aneesh Durai, John DeNero, J. D. Zamfirescu-Pereira, Narges Norouzi |
SIGCSE (2) | 5 |
| 2025 | Spotting AI Missteps: Students Take on LLM Errors in CS1
Samantha Boatright Smith, Heather Wei, Abigail O'Neill, Aneesh Durai, John DeNero, J. D. Zamfirescu-Pereira, Narges Norouzi |
SIGCSE (2) | 6 |
| 2025 | Pensieve Discuss: Scalable Small-Group CS Tutoring System with AI
Yoonseok Yang, Jack Liu, J. D. Zamfirescu-Pereira, John DeNero |
SIGCSE (2) | 3 |
| 2025 | 61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?)abstractLLM-based chatbots enable students to get immediate, interactive help on homework assignments, but even a thoughtfully-designed bot may not serve all pedagogical goals. We report here on the development and deployment of a GPT-4-based interactive homework assistant ("61A Bot'') for students in a large CS1 course; over 2000 students made over 100,000 requests of our Bot across two semesters. Our assistant offers one-shot, contextual feedback within the command-line "autograder'' students use to test their code. Our Bot wraps student code in a custom prompt that supports our pedagogical goals and avoids providing solutions directly. Analyzing student feedback, questions, and autograder data, we find reductions in homework-related question rates in our course forum, as well as reductions in homework completion time when our Bot is available. For students in the 50th -80th percentile, reductions can exceed 30 minutes per assignment, up to 50% less time than students at the same percentile rank in prior semesters. Finally, we discuss these observations, potential impacts on student learning, and other potential costs and benefits of AI assistance in CS1. J. D. Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, Narges Norouzi |
SIGCSE (1) | 1 |
| 2024 | Prompting for Discovery: Flexible Sense-Making for AI Art-Making with DreamsheetsabstractDesign space exploration (DSE) for Text-to-Image (TTI) models entails navigating a vast, opaque space of possible image outputs, through a commensurately vast input space of hyperparameters and prompt text. Perceptually small movements in prompt-space can surface unexpectedly disparate images. How can interfaces support end-users in reliably steering prompt-space explorations towards interesting results? Our design probe, DreamSheets, supports user-composed exploration strategies with LLM-assisted prompt construction and large-scale simultaneous display of generated results, hosted in a spreadsheet interface. Two studies, a preliminary lab study and an extended two-week study where five expert artists developed custom TTI sheet-systems, reveal various strategies for targeted TTI design space exploration—such as using templated text generation to define and layer semantic “axes” for exploration. We identified patterns in exploratory structures across our participants’ sheet-systems: configurable exploration “units” that we distill into a UI mockup, and generalizable UI components to guide future interfaces. Shm Garanganao Almeda, J. D. Zamfirescu-Pereira, Kyu Won Kim, Pradeep Mani Rathnam, Björn Hartmann |
CHI | 2 |
| 2024 | Rambler: Supporting Writing With Speech via LLM-Assisted Gist ManipulationabstractDictation enables efficient text input on mobile devices. However, writing with speech can produce disfluent, wordy, and incoherent text and thus requires heavy post-processing. This paper presents Rambler, an LLM-powered graphical user interface that supports gist-level manipulation of dictated text with two main sets of functions: gist extraction and macro revision. Gist extraction generates keywords and summaries as anchors to support the review and interaction with spoken text. LLM-assisted macro revisions allow users to respeak, split, merge, and transform dictated text without specifying precise editing locations. Together they pave the way for interactive dictation and revision that help close gaps between spontaneously spoken words and well-structured writing. In a comparative study with 12 participants performing verbal composition tasks, Rambler outperformed the baseline of a speech-to-text editor + ChatGPT, as it better facilitates iterative revisions with enhanced user control over the content while supporting surprisingly diverse user strategies. Susan Lin, Jeremy Warner, J. D. Zamfirescu-Pereira, Matthew G. Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Björn Hartmann, Can Liu 0003 |
CHI | 3 |
| 2024 | Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesabstractDue to the cumbersome nature of human evaluation and limitations of code-based evaluation, Large Language Models (LLMs) are increasingly being used to assist humans in evaluating LLM outputs. Yet LLM-generated evaluators simply inherit all the problems of the LLMs they evaluate, requiring further human validation. We present a mixed-initiative approach to “validate the validators”—aligning LLM-generated evaluation functions (be it prompts or code) with human requirements. Our interface, EvalGen, provides automated assistance to users in generating evaluation criteria and implementing assertions. While generating candidate implementations (Python functions, LLM grader prompts), EvalGen asks humans to grade a subset of LLM outputs; this feedback is used to select implementations that better align with user grades. A qualitative study finds overall support for EvalGen but underscores the subjectivity and iterative nature of alignment. In particular, we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria. What is more, some criteria appear dependent on the specific LLM outputs observed (rather than independent and definable a priori), raising serious questions for approaches that assume the independence of evaluation from observation of model outputs. We present our interface and implementation details, a comparison of our algorithm with a baseline approach, and implications for the design of future LLM evaluation assistants. Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran, Ian Arawjo |
UIST | 2 |
| 2023 | Herding AI Cats: Lessons from Designing a Chatbot by Prompting GPT-3abstractPrompting Large Language Models (LLMs) is an exciting new approach to designing chatbots. But can it improve LLM’s user experience (UX) reliably enough to power chatbot products? Our attempt to design a robust chatbot by prompting GPT-3/4 alone suggests: not yet. Prompts made achieving “80%” UX goals easy, but not the remaining 20%. Fixing the few remaining interaction breakdowns resembled herding cats: We could not address one UX issue or test one design solution at a time; instead, we had to handle everything everywhere all at once. Moreover, because no prompt could make GPT reliably say “I don’t know” when it should, the user-GPT conversations had no guardrails after a breakdown occurred, often leading to UX downward spirals. These risks incentivized us to design highly prescriptive prompts and scripted bots, counter to the promises of LLM-powered chatbots. This paper describes this case study, unpacks prompting’s fickleness and its impact on UX design processes, and discusses implications for LLM-based design methods and tools. J. D. Zamfirescu-Pereira, Heather Wei, Amy Xiao, Kitty Gu, Grace Jung, Matthew G. Lee, Björn Hartmann, Qian Yang 0004 |
Conference on Designing Interactive Systems | 1 |
| 2023 | Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsabstractPre-trained large language models (“LLMs”) like GPT-3 can engage in fluent, multi-turn instruction-taking out-of-the-box, making them attractive materials for designing natural language interactions. Using natural language to steer LLM outputs (“prompting”) has emerged as an important design technique potentially accessible to non-AI-experts. Crafting effective prompts can be challenging, however, and prompt-based interactions are brittle. Here, we explore whether non-AI-experts can successfully engage in “end-user prompt engineering” using a design probe—a prototype LLM-based chatbot design tool supporting development and systematic evaluation of prompting strategies. Ultimately, our probe participants explored prompt designs opportunistically, not systematically, and struggled in ways echoing end-user programming systems and interactive machine learning systems. Expectations stemming from human-to-human instructional experiences, and a tendency to overgeneralize, were barriers to effective prompt design. These findings have implications for non-AI-expert-facing LLM-based tool design and for improving LLM-and-prompt literacy among programmers and the public, and present opportunities for further research. J. D. Zamfirescu-Pereira, Richmond Y. Wong, Björn Hartmann, Qian Yang 0004 |
CHI | 1 |
| 2019 | How People Experience Autonomous Intersections: Taking a First-Person PerspectiveabstractTop-down simulations of autonomous intersections neglect considerations for the human experience of being in cars driving through these autonomous intersections. To understand the impact that perspective has on perception of autonomous intersections, we conducted a driving simulator experiment and studied the experience in terms of perception, feelings, and pleasure. Based on this data, we discuss experiential factors of autonomous intersections that are perceived as beneficial or detrimental for the future driver. Furthermore, we present what the change of perspective implies for designing intersection models, future in-car interfaces and simulation techniques. Sven Krome, David Goedicke, Thomas J. Matarazzo, Zimeng Zhu, J. D. Zamfirescu-Pereira, Wendy Ju |
AutomotiveUI | 6 |
| 2019 | Heimdall: A Remotely Controlled Inspection Workbench For Debugging Microcontroller ProjectsabstractStudents and hobbyists build embedded systems that combine sensing, actuation and microcontrollers on solderless breadboards. To help students debug such circuits, experienced teachers apply visual inspection, targeted measurements, and circuit modifications to diagnose and localize the problem(s). However, experienced helpers may not always be available to review student projects in person. To enable remote debugging of circuit problems, we introduce Heimdall, a remote electronics workbench that allows experts to visually inspect a student's circuit; perform measurements; and to re-wire and inject test signals. These interactions are enabled by an actuated inspection camera; an augmented breadboard that enables flexible configuration of row connectivity and measurement/injection lines; and a web-based UI that teachers can use to perform measurements through interaction with the captured images. We demonstrate that common issues arising in embedded electronics classes can be successfully diagnosed remotely and report on preliminary user feedback from teaching assistants who frequently debug circuits. Mitchell Karchemsky, J. D. Zamfirescu-Pereira, Kuan-Ju Wu, François Guimbretière, Björn Hartmann |
CHI | 2 |