VLDB 2026 Research / reviewers in the wild / expert
Ian Arawjo
dblp:194/9606 · also Ian A. Arawjo
· DBLP profile ↗
19ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0001-8910-0822ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Metaphors for Memory: Charting a Design Space of AI Memory Tools and InterfacesabstractAI memory is becoming central to AI systems that aim to support personal and professional work. Yet, designing interfaces for AI memory is not well understood. How should we design for user interaction with AI memories—for instance, what kinds of operations might users want to perform on memory, either now or in the future? We survey the design of tools and interfaces around “AI memory” to chart a design space of common design patterns, architectures, and operations, and identify dominant metaphors and gaps in the space. Then, we apply generative metaphorical design to expand the design space for AI memory, exploring less-dominant metaphors of software version control, Zettelkasten, requirements, personal diaries, community archives, cultural probes, and science fiction. Our work offers rich opportunities for gaps and emergent needs that future interfaces for AI memory might address. Munyeong Kim, Michalis Famelis, Ian Arawjo |
DIS | 3 |
| 2026 | Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and ConsiderationsabstractWhat should HCI scholars consider when reporting and reviewing papers that involve LLM-integrated systems? We interview 18 authors of LLM-integrated system papers on their authoring and reviewing experiences. We find that norms of trust-building between authors and reviewers appear to be eroded by the uncertainty of LLM behavior and hyperbolic rhetoric surrounding AI. Authors perceive that reviewers apply uniquely skeptical and inconsistent standards towards papers that report LLM-integrated systems, and mitigate mistrust by adding technical evaluations, justifying usage, and de-emphasizing LLM presence. Authors’ views challenge blanket directives to report all prompts and use open models, arguing that prompt reporting is context-dependent and that proprietary model usage can be justified despite ethical concerns. Finally, some tensions in peer review appear to stem from clashes between the norms and values of HCI and ML/NLP communities, particularly around what constitutes a contribution and an appropriate level of technical rigor. Based on our findings and additional feedback from six expert HCI researchers, we present a set of considerations for authors, reviewers, and HCI communities around reporting and reviewing papers that involve LLM-integrated systems. Karla Felix Navarro, Eugene Syriani, Ian Arawjo |
CHI | 3 |
| 2026 | How Notations Evolve: A Historical Analysis with Implications for Supporting User-Defined AbstractionsabstractTraditional human-computer interaction takes place through formally-specified systems like structured UIs and programming languages. Recent AI systems promise a new set of informal interactions with computers through natural language and other notational forms. These informal interactions can then lead to formal representations, but depend upon pre-existing formalisms known to both humans and AI. What about novel formalisms and notations? How are new abstractions created, evolved, and incrementally formalized over time—and how might new systems, in turn, be explicitly designed to support these processes? We conduct a comparative historical analysis of notation development to identify some relevant characteristics. These include three social stages of notation development: invention & incubation, dispersion & divergence, and institutionalization & sanctification, as well as three functional stages: descriptive, generative, and evaluative. Within and across these stages, we detail several patterns, such as the role of linking and grounding metaphors, dimensions of meaningful variation, and analogical alignment. Finally, we offer some implications for design. Jingyue Zhang, J. D. Zamfirescu-Pereira, Elena L. Glassman, Damien Masson, Ian Arawjo |
CHI | 5 |
| 2025 | Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming SupportabstractAI programming tools enable powerful code generation, and recent prototypes attempt to reduce user effort with proactive AI agents, but their impact on programming workflows remains unexplored. We introduce and evaluate Codellaborator, a design probe LLM agent that initiates programming assistance based on editor activities and task context. We explored three interface variants to assess trade-offs between increasingly salient AI support: prompt-only, proactive agent, and proactive agent with presence and context (Codellaborator). In a within-subject study (N=18), we find that proactive agents increase efficiency compared to prompt-only paradigm, but also incur workflow disruptions. However, presence indicators and interaction context support alleviated disruptions and improved users' awareness of AI processes. We underscore trade-offs of Codellaborator on user control, ownership, and code understanding, emphasizing the need to adapt proactivity to programming processes. Our research contributes to the design exploration and evaluation of proactive AI systems, presenting design implications on AI-integrated programming workflow. Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia, Ziang Xiao, Tovi Grossman, Yan Chen 0033 |
CHI | 3 |
| 2025 | ChainBuddy: An AI-assisted Agent System for Generating LLM Pipelines
Jingyue Zhang, Ian Arawjo |
CHI | 2 |
| 2025 | Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale
Priyan Vaithilingam, Munyeong Kim, Frida-Cecilia Acosta-Parenteau, Amine Mhedhbi, Elena L. Glassman, Ian Arawjo |
UIST | 7 |
| 2025 | Democratizing Game Modding with GenAI: A Case Study of StarCharM, a Stardew Valley Character MakerabstractGame modding offers unique and personalized gaming experiences, but the technical complexity of creating mods often limits participation to skilled users. We envision a future where every player can create personalized mods for their games. To explore this space, we designed StarCharM, a GenAI-based non-player character (NPC) creator for Stardew Valley. Our tool enables players to iteratively create new NPC mods, requiring minimal user input while allowing for fine-grained adjustments through user control. We conducted a user study with ten Stardew Valley players who had varied mod usage experiences to understand the impacts of StarCharM and provide insights into how GenAI tools may reshape modding, particularly in NPC creation. Participants expressed excitement in bringing their character ideas to life, although they noted challenges in generating rich content to fulfill complex visions. While they believed GenAI tools like StarCharM can foster a more diverse modding community, some voiced concerns about diminished originality and community engagement that may come with such technology. Our findings provided implications and guidelines for the future of GenAI-powered modding tools and co-creative modding practices. Hamid Zand Miralvand, Mohammad Ronagh Nikghalb, Mohammad Darandeh, Abidullah Khan, Ian Arawjo, Jinghui Cheng 0001 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Imagining a Future of Designing with AI: Dynamic Grounding, Constructive Negotiation, and Sustainable MotivationabstractWe ideate a future design workflow that involves AI technology. Drawing from activity and communication theory, we attempt to isolate the new value that large AI models can provide design compared to past technologies. We arrive at three affordances—dynamic grounding, constructive negotiation, and sustainable motivation—that summarize latent qualities of natural language-enabled foundation models that, if explicitly designed for, can support the process of design. Through design fiction, we then imagine a future interface as a diegetic prototype, the story of Squirrel Game, that demonstrates each of our three affordances in a realistic usage scenario. Our design process, terminology, and diagrams aim to contribute to future discussions about the relative affordances of AI technology with regard to collaborating with human designers. Priyan Vaithilingam, Ian Arawjo, Elena L. Glassman |
Conference on Designing Interactive Systems | 2 |
| 2024 | ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingabstractEvaluating outputs of large language models (LLMs) is challenging, requiring making—and making sense of—many responses. Yet tools that go beyond basic prompting tend to require knowledge of programming APIs, focus on narrow domains, or are closed-source. We present ChainForge, an open-source visual toolkit for prompt engineering and on-demand hypothesis testing of text generation LLMs. ChainForge provides a graphical interface for comparison of responses across models and prompt variations. Our system was designed to support three tasks: model selection, prompt template design, and hypothesis testing (e.g., auditing). We released ChainForge early in its development and iterated on its design with academics and online users. Through in-lab and interview studies, we find that a range of people could use ChainForge to investigate hypotheses that matter to them, including in real-world settings. We identify three modes of prompt engineering and LLM hypothesis testing: opportunistic exploration, limited evaluation, and iterative refinement. Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, Elena L. Glassman |
CHI | 1 |
| 2024 | An AI-Resilient Text Rendering Technique for Reading and Skimming DocumentsabstractReaders find text difficult to consume for many reasons. Summarization can address some of these difficulties, but introduce others, such as omitting, misrepresenting, or hallucinating information, which can be hard for a reader to notice. One approach to addressing this problem is to instead modify how the original text is rendered to make important information more salient. We introduce Grammar-Preserving Text Saliency Modulation (GP-TSM), a text rendering method with a novel means of identifying what to de-emphasize. Specifically, GP-TSM uses a recursive sentence compression method to identify successive levels of detail beyond the core meaning of a passage, which are de-emphasized by rendering words in successively lighter but still legible gray text. In a lab study (n=18), participants preferred GP-TSM over pre-existing word-level text rendering methods and were able to answer GRE reading comprehension questions more efficiently. Ziwei Gu, Ian Arawjo, Kenneth Li 0002, Jonathan K. Kummerfeld, Elena L. Glassman |
CHI | 2 |
| 2024 | Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesabstractDue to the cumbersome nature of human evaluation and limitations of code-based evaluation, Large Language Models (LLMs) are increasingly being used to assist humans in evaluating LLM outputs. Yet LLM-generated evaluators simply inherit all the problems of the LLMs they evaluate, requiring further human validation. We present a mixed-initiative approach to “validate the validators”—aligning LLM-generated evaluation functions (be it prompts or code) with human requirements. Our interface, EvalGen, provides automated assistance to users in generating evaluation criteria and implementing assertions. While generating candidate implementations (Python functions, LLM grader prompts), EvalGen asks humans to grade a subset of LLM outputs; this feedback is used to select implementations that better align with user grades. A qualitative study finds overall support for EvalGen but underscores the subjectivity and iterative nature of alignment. In particular, we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria. What is more, some criteria appear dependent on the specific LLM outputs observed (rather than independent and definable a priori), raising serious questions for approaches that assume the independence of evaluation from observation of model outputs. We present our interface and implementation details, a comparison of our algorithm with a baseline approach, and implications for the design of future LLM evaluation assistants. Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran, Ian Arawjo |
UIST | 5 |
| 2022 | Notational Programming for Notebook Environments: A Case Study with Quantum CircuitsabstractWe articulate a vision for computer programming that includes pen-based computing, a paradigm we term notational programming. Notational programming blurs contexts: certain typewritten variables can be referenced in handwritten notation and vice-versa. To illustrate this paradigm, we developed an extension, Notate, to computational notebooks which allows users to open drawing canvases within lines of code. As a case study, we explore quantum programming and designed a notation, Qaw, that extends quantum circuit notation with abstraction features, such as variable-sized wire bundles and recursion. Results from a usability study with novices suggest that users find our core interaction of implicit cross-context references intuitive, but suggests further improvements to debugging infrastructure, interface design, and recognition rates. Throughout, we discuss questions raised by the notational paradigm, including a shift from ‘recognition’ of notations to ‘reconfiguration’ of practices and values around programming, and from ‘sketching’ to writing and drawing, or what we call ‘notating.’ Ian Arawjo, Anthony DeArmas, Michael Roberts, Shrutarshi Basu, Tapan S. Parikh |
UIST | 1 |
| 2021 | Intercultural Computing Education: Toward Justice Across DifferenceabstractEven in the turn toward justice-oriented pedagogy, computing education tends to overlook the quality of intergroup relationships, which risks entrenching division. In this article, we establish an intercultural approach to computing education, informed by intercultural and peace education, prejudice reduction, and the sociology of racism and ethnicity. We outline three major concerns of intercultural computing: shifting from content toward relationships, from cultural responsiveness to cultural reflexivity, and from identity to identification. For the last, we complicate discourses of race and identity widespread in U.S. education. Drawing from studies of youth programming classes in East Africa and U.S. contexts, we then reflect on our attempts to address the first shift of fostering relationships across difference. We highlight three promising design tactics: intergroup pairing, interdependent programming, and making relational goals explicit. Overall, we find that computing can indeed be a site of intergroup bonding across difference, but that bonding can carry complications and tensions with other equity goals and tactics. Rather than framing justice-oriented CS primarily as changes to the aims of computational learning, we argue that future work should explore making relational goals explicit and teach students how to attend to friction. Ian Arawjo, Ariam Mogos |
ACM Trans. Comput. Educ. | 1 |
| 2020 | To Write Code: The Cultural Fabrication of Programming Notation and PracticeabstractWriting and its means have become detached. Unlike written and drawn practices developed prior to the 20th century, notation for programming computers developed in concert and conflict with discretizing infrastructure such as the shift-key typewriter and data processing pipelines. In this paper, I recall the emergence of high-level notation for representing computation. I show how the earliest inventors of programming notations borrowed from various written cultural practices, some of which came into conflict with the constraints of digitizing machines, most prominently the typewriter. As such, I trace how practices of "writing code" were fabricated along social, cultural, and material lines at the time of their emergence. By juxtaposing early visions with the modern status quo, I question long-standing terminology, dichotomies, and epistemological tendencies in the field of computer programming. Finally, I argue that translation work is a fundamental property of the practice of writing code by advancing an intercultural lens on programming practice rooted in history. Ian Arawjo |
CHI | 1 |
| 2020 | Exploring Intercultural Approaches to Resolving Sociocultural Tension in CS ClassesabstractCS education currently considers problems of access or participation for various marginalized groups, often in the U.S.-based settings. Yet even where such thorny problems are resolved, equity research suggests that sociocultural tension in classrooms -often between students themselves -may reproduce disparities and boundaries of the wider society, even (or especially) in politically- or ethically-conscious curricula (as argued recently by STEM equity researchers such as Sepehr Vahil and Na'ilah Nasir). A challenge thus remains in some diverse classrooms and contexts to teach understanding between all students, rather than purely computing. This lightning talk introduces our effort for intercultural computing education called the Nairobi Play Project, a computational thinking (CT) course for multi-ethnic East African youth that integrates intercultural competence and peace-building activities. We talk about our current exploration of theoretical and pedagogical alignments (and tensions) between CT and intercultural learning, and our current challenges in dealing with the complications of intercultural computing endeavors. We hope the talk opens up further discussion and debate on the role of CS educational spaces as sites sociocultural learning and disruption of the reproduction of disparities. Ian Arawjo, Ariam Mogos |
SIGCSE | 1 |
| 2019 | Computing Education for Intercultural Learning: Lessons from the Nairobi Play ProjectabstractThis paper explores computing education as a potential site for intercultural learning and encounter in post-conflict environments. It reports on ethnographic fieldwork from the Nairobi Play Project, a constructionist educational program serving adolescents aged 14-18 in urban and rural multi-ethnic refugee communities in Kenya. While the program offers programming and game design instruction, an equal goal is to foster interaction, collaboration, dialogue and understanding across cultural backgrounds. Based on fieldwork from two project cycles involving 5 after-school classes of 12-24 students each, we describe key affordances for encounter, important resistances to be managed or overcome, and emergent complications in the execution of such programs. We argue that many important accomplishments of intercultural learning occur through moments of friction, breakdowns, and gaps -- for example, technical challenges that produce sites of shared humour; frictions between intercultural activities and computing activities; acts of disrupting order; and unstructured time that students collaboratively fill in. We also describe significant complications in such programs, including pressures to adopt norms and practices consistent with dominant or majority cultures, and instances of intercultural bonding over artefacts with xenophobic themes. We reflect on the implications of these phenomena for the design of future programs that use computing as a backbone for intercultural learning or diversity and inclusion efforts in CSCW, ICTD, and allied fields of work. Ian Arawjo, Ariam Mogos, Steven J. Jackson, Tapan S. Parikh, Kentaro Toyama |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2017 | Teaching Programming with Gamified SemanticsabstractDominant approaches to programming education emphasize program construction over language comprehension. We present Reduct, an educational game embodying a new, comprehension-first approach to teaching novices core programming concepts which include functions, Booleans, equality, conditionals, and mapping functions over sets. In this novel teaching strategy, the player executes code using reduction-based operational semantics. During gameplay, code representations fade from concrete, block-based graphics to the actual syntax of JavaScript ES2015. We describe our design rationale and report on the results of a study evaluating the efficacy of our approach on young adults (18+) without prior coding experience. In a short timeframe, novices demonstrated promising learning of core concepts expressed in actual JavaScript. We also present results from an online deployment. Finally, we discuss ramifications for the design of future computational thinking games. Ian Arawjo, Cheng-Yao Wang, Andrew C. Myers, Erik Andersen 0001, François Guimbretière |
CHI | 1 |
| 2017 | TypeTalker: A Speech Synthesis-Based Multi-Modal Commenting SystemabstractSpeech commenting systems have been shown to facilitate asynchronous online communication from educational discussion to writing feedback. However, the production of speech comments introduces several challenges to users, including overcoming self-consciousness and time consuming editing. In this paper, we introduce TypeTalker, a speech commenting interface that presents speech as a synthesized generic voice to reduce speaker self-consciousness, while retaining the expressivity of the original speech with natural breaks and co-expressive gestures. TypeTalker streamlines speech editing through a simple textbox that respects temporal alignment across edits. A comparative evaluation shows that TypeTalker reduces speech anxiety during live-recording, and offers easier and more effective speech editing facilities than the previous state-of-the-art interface technique. A follow-up study on recipient perceptions of the produced comments suggests that while TypeTalker's generic voice may be traded-off with a loss of personal touch, it can also enhance the clarity of speech by refining the original speech's speed and accent. Ian Arawjo, Dongwook Yoon, François Guimbretière |
CSCW | 1 |
| 2017 | Distraction or Life Saver?: The Role of Technology in Undergraduate Students' Boundary Management StrategiesabstractPrevious research has shown that communication technologies may make it challenging for working professionals to manage the boundaries between their work life and home life. For college students, however, there is a less clear definition of what constitutes work and what constitutes home life. As a result, students may use different boundary management strategies than working professionals. To explore this issue, we interviewed 29 undergraduates about how they managed boundaries between different areas of their life. Interviewees reported maintaining flexible and permeable boundaries that are not bounded physically or temporally. They used both technological and non-technological strategies to manage different life spheres. Interviewees saw technology as a major source of boundary violations but also as a boundary managing strategy that allowed them to achieve better life balance. Based on these findings, we propose design implications for tools to better support the boundary management processes of undergraduate students. Hajin Lim, Ian Arawjo, Yaxian Xie, Negar Khojasteh, Susan R. Fussell |
Proc. ACM Hum. Comput. Interact. | 2 |