VLDB 2026 Research / reviewers in the wild / expert
Toby Jia-Jun Li
dblp:158/9197
· DBLP profile ↗
65ranked-venue papers
9as first author
53since 2021 · last 2026
0000-0001-7902-7625ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 47 · 7 first-author · 38 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PrivacyMotiv: Vulnerability-Centered Persona Journeys for Empathic Privacy Reviews in UX DesignabstractUX professionals routinely conduct design reviews, yet privacy concerns are often overlooked, not only due to limited tools, but more fundamentally from low intrinsic motivation, driven by limited privacy knowledge, weak empathy for unexpectedly affected users, and low autonomy in identifying harms. We present PrivacyMotiv, an LLM-powered system that generates vulnerability-centered personas, persona journey stories, and traceable design diagnoses grounded in lo-fi user flows to support privacy-oriented UX design review. In a within-subjects study with professional UX practitioners (N=16), PrivacyMotiv significantly improved empathy, intrinsic motivation, and perceived usefulness, with participants identifying 59% more privacy issues and proposing 70% more redesign solutions compared to self-proposed methods. This work contributes empirical insight into motivational barriers in privacy-aware UX and a structured, narrative-driven approach for integrating privacy review into early-stage UX practice. Zeya Chen, Jianing Wen, Yaxing Yao, Toby Jia-Jun Li, Tianshi Li 0001 |
DIS | 4 |
| 2026 | OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior SimulationabstractZiyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini, Bo Sun, Yakov Bart, Weimin Lyu, Jiri Gesi, Tian Wang, Jing Huang, Yu Su, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia Chilton, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxuan Lu 0003, Amirali Amini, Yakov Bart, Weimin Lyu, Jiri Gesi, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia B. Chilton, Dakuo Wang |
ACL (1) | 14 |
| 2026 | EyeMulator: Improving Code Language Models by Mimicking Human Visual AttentionabstractYifan Zhang, Chen Huang, Yueke Zhang, Jiahao Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yifan Zhang 0013, Chen Huang 0006, Yueke Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang 0015 |
ACL (1) | 5 |
| 2026 | Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-EditingabstractHumans think visually—we remember in images, dream in pictures, and use visual metaphors to communicate. Yet, most creative writing tools remain text-centric, limiting how writers plan and translate ideas. We present Vistoria, a system for synchronized image-text co-editing in fictional story writing. A formative Wizard-of-Oz co-design study with 10 story writers revealed how sketches, images, and text serve as essential elements for ideation and organization. Drawing on theories of Instrumental Interaction, Vistoria introduces instrumental operations—Lasso, Collage, Perspective Shift, and Filter that enable seamless narrative exploration across modalities. A controlled study with 12 participants shows that co-editing enhances expressiveness, immersion, and collaboration, opening space for writers to follow divergent story directions and craft more vivid, detailed narratives. While multimodality increased cognitive demand, participants reported stronger senses of ownership and agency. These findings demonstrate how multimodal co-editing expands creative potential by balancing abstraction and concreteness in narrative development. Kexue Fu 0002, Jingfei Huang, Long Ling, Sumin Hong 0001, Yihang Zuo, Ray LC, Toby Jia-Jun Li |
CHI | 7 |
| 2026 | Crepe: A Mobile Screen Data Collector Using Graph QueryabstractCollecting mobile datasets remains challenging for academic researchers due to limited data access and technical barriers. Commercial organizations often possess exclusive access to mobile data, leading to a "data monopoly" that restricts the independence of academic research. Existing open-source mobile data collection frameworks primarily focus on mobile sensing data rather than screen content, which is crucial for various research studies. We present Crepe, a no-code Android app that enables researchers to collect information displayed on screen through simple demonstrations of target data. Crepe utilizes a novel Graph Query technique which augments the structures of mobile UI screens to support flexible identification, location, and collection of specific data pieces. The tool emphasizes participants' privacy and agency by providing full transparency over collected data and allowing easy opt-out. We designed and built Crepe for research purposes only and in scenarios where researchers obtain explicit consent from participants. Code for Crepe will be open-sourced to support future academic research data collection. Yuwen Lu, Meng Chen 0020, Victor V. Cox, Yang Yang 0008, Meng Jiang 0001, Jay Brockman, Tamara Kay, Toby Jia-Jun Li |
CHI | 9 |
| 2026 | Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation CriteriaabstractLarge Language Models (LLMs) are increasingly utilized for domain-specific tasks, yet evaluating their outputs remains challenging. A common strategy is to apply evaluation criteria to assess alignment with domain-specific standards, yet little is understood about how criteria differ across sources or where each type is most useful in the evaluation process. This study investigates criteria developed by domain experts, lay users, and LLMs to identify their complementary roles within an evaluation workflow. Results show that experts produce fact-based criteria with long-term value, lay users emphasize usability with a shorter-term focus, and LLMs target procedural checks for immediate task requirements. We also examine how criteria evolve between a priori and a posteriori phases, noting drift across stages as well as convergence in the a posteriori phase. Based on our observations, we propose design guidelines for a staged evaluation workflow combining the complementary strengths of these sources to balance quality, cost, and scalability. Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A. Metoyer, Toby Jia-Jun Li |
CHI | 5 |
| 2026 | Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human OversightabstractThe dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks from high-level intents, understanding how dark patterns affect agents is increasingly important. We present a two-phase empirical study examining how agents, human participants, and human-AI teams respond to 16 types of dark patterns across diverse scenarios. Phase 1 highlights that agents often fail to recognize dark patterns, and even when aware, prioritize task completion over protective action. Phase 2 revealed divergent failure modes: humans succumb due to cognitive shortcuts and habitual compliance, while agents falter from procedural blind spots. Human oversight improved avoidance but introduced costs such as attentional tunneling and cognitive load. Our findings show neither humans nor agents are uniformly resilient, and collaboration introduces new vulnerabilities, suggesting design needs for transparency, adjustable autonomy, and oversight. Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye 0001, Tianshi Li 0001, Ziang Xiao, Yaxing Yao, Toby Jia-Jun Li |
CHI | 14 |
| 2026 | Through the Lens of Human-Human Collaboration: An Configurable Research Platform for Exploring Human-Agent CollaborationabstractIntelligent systems have traditionally been designed as tools rather than collaborators, often lacking critical characteristics that collaboration partnerships require. Recent advances in large language model (LLM) agents open new opportunities for human-LLM-agent collaboration by enabling natural communication and various social and cognitive behaviors. Yet it remains unclear whether principles of computer-mediated collaboration established in HCI and CSCW persist, change, or fail when humans collaborate with LLM agents. To support systematic investigations of these questions, we introduce an open and configurable research platform for HCI researchers1. The platform’s modular design allows seamless adaptation of classic CSCW experiments and manipulation of theory-grounded interaction controls. We demonstrate the platform’s research efficacy and usability through three case studies: (1) two Shape FactoryHidden Profile experiment for information pooling with 16 participants, and (3) a participatory cognitive walkthrough with five HCI researchers to refine workflows of researcher interface for experiment setup and analysis. Bingsheng Yao, Jiaju Chen, April Yi Wang, Toby Jia-Jun Li, Dakuo Wang |
CHI | 5 |
| 2026 | My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber FandomabstractAI VTubers, where the performer is not human but algorithmically generated, introduce a new context for fandom. While human VTubers have been substantially studied for their cultural appeal, parasocial dynamics, and community economies, little is known about how audiences engage with their AI counterparts. To address this gap, we present a qualitative study of Neuro-sama, the most prominent AI VTuber. Our findings show that engagement is anchored in active co-creation: audiences are drawn by the AI’s unpredictable yet entertaining interactions, cement loyalty through collective emotional events that trigger anthropomorphic projection, and sustain attachment via the AI’s consistent persona. Financial support emerges not as a reward for performance but as a participatory mechanism for shaping livestream content, establishing a resilient fan economy built on ongoing interaction. These dynamics reveal how AI Vtuber fandom reshapes fan–creator relationships and offer implications for designing transparent and sustainable AI-mediated communities. Jiayi Ye, Yue Huang 0001, Yanfang Ye 0001, Toby Jia-Jun Li, Xiangliang Zhang 0001 |
CHI | 5 |
| 2026 | The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction OutcomesabstractLarge Language Model (LLM)-powered web GUI agents are increasingly automating everyday online tasks. Despite their popularity, little is known about how users’ preferences and values impact agents’ reasoning and behavior. In this work, we investigate how both explicit and implicit user preferences, as well as the underlying user values, influence agent decision-making and action trajectories. We built a controlled testbed of 14 common interactive web tasks, spanning shopping, travel, dining, and housing, each replicated from real websites and integrated with a low-fidelity LLM-based recommender system. We injected 12 human preferences and values as personas into four state-of-the-art agents and systematically analyzed their task behaviors. Our results show that preference and value-infused prompts consistently guided agents toward outcomes that reflected these preferences and values. While the absence of user preference or value guidance led agents to exhibit a strong efficiency bias and employ shortest-path strategies, their presence steered agents’ behavior trajectories through the greater use of corresponding filters and interactive web features. Despite their influence, dominant interface cues, such as discounts and advertisements, frequently overrode these effects, shortening the agents’ action trajectories and inducing rationalizations that masked rather than reflected value-consistent reasoning. The contributions of this paper are twofold: (1) an open-source testbed for studying the influence of values in agent behaviors, and (2) an empirical investigation of how user preferences and values shape web agent behaviors. Simret Araya Gebreegziabher, Yukun Yang 0008, Charles Chiang, Hojun Yoo, Hyo Jin Do, Zahra Ashktorab, Werner Geyer, Diego Gómez-Zará, Toby Jia-Jun Li |
IUI | 10 |
| 2026 | Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic StudyabstractLarge Language Models (LLMs) are increasingly developed for use in complex professional domains, yet little is known about how teams design and evaluate these systems in practice. This paper examines the challenges and trade-offs in LLM development through a 12-week ethnographic study of a team building a pedagogical chatbot. The researcher observed design and evaluation activities and conducted interviews with both developers and domain experts. Analysis revealed four key practices: creating workarounds for data collection, turning to augmentation when expert input was limited, co-developing evaluation criteria with experts, and adopting hybrid expert–developer–LLM evaluation strategies. These practices show how teams made strategic decisions under constraints and demonstrate the central role of domain expertise in shaping the system. Challenges included expert motivation and trust, difficulties structuring participatory design, and questions around ownership and integration of expert knowledge. We propose design opportunities for future LLM development workflows that emphasize AI literacy, transparent consent, and frameworks recognizing evolving expert roles. Annalisa Szymanski, Oghenemaro Anuyah, Toby Jia-Jun Li, Ronald A. Metoyer |
IUI | 3 |
| 2026 | A comprehensive survey of AI agents in healthcareabstractOBJECTIVE: This survey aims to systematically map the rapidly evolving landscape of AI agents in healthcare. It addresses the critical need to adapt general-purpose agentic frameworks characterized by autonomy, planning, and tool use to the high-stakes, safety-critical constraints of medical decision-making and patient care. METHODS: We conducted a comprehensive review of over 200 recent studies, synthesizing literature from major academic databases. We developed a holistic taxonomy that traces the full lifecycle of healthcare agents, analyzing perception modalities, core technical architectures, and evaluation protocols specific to autonomous systems. RESULTS: The review presents a quantitative landscape analysis showing exponential growth in the field. We structure the domain into three pillars: (1) Perception of multi-modal clinical data (e.g., EHR, imaging, genomics); (2) Agent Capabilities, including tool use, reasoning, memory, and multi-agent collaboration; and (3) an Application Ecosystem organized by stakeholder roles (clinicians, patients, researchers, and administrators). Additionally, we categorize evaluation frameworks, and discuss the deployment readiness of current systems across technical, evidentiary, and governance dimensions. Finally, we identify challenges for advancing healthcare agents from controlled evaluation toward real-world clinical integration. A continuously updated repository of related papers is available at https://github.com/AgenticHealthAI/Awesome-AI-Agents-for-Healthcare. CONCLUSION: AI agents offer significant potential to enhance healthcare through autonomous reasoning and workflow integration. However, current research remains largely concentrated in benchmark and controlled evaluation settings, and the translation into clinical practice will require advances in reliability, privacy protection, governance, and operational integration. Gelei Xu, Yixiong Chen, Yuying Duan, Shuqing Wu, Haoxinran Yu, Ching-Hao Chiu, Juntong Ni, Ningzhi Tang, Toby Jia-Jun Li, Alan L. Yuille, Wei Jin 0009, Yiyu Shi 0001 |
J. Biomed. Informatics | 10 |
| 2026 | Context-aware code summary generation
Chia-Yi Su, Aakash Bansal, Yu Huang 0015, Toby Jia-Jun Li, Collin McMillan |
J. Syst. Softw. | 4 |
| 2026 | Investigating the Feasibility of Conducting Webcam-Based Eye-Tracking Studies in Code ComprehensionabstractResearchers in Software Engineering (SE) often use onsite screen-mounted eye-tracking experiments to investigate programmers’ visual attention patterns in various programming activities. The pandemic and the difficulty of recruiting many participants, especially those with special expertise in SE, have hastened the shift towards conducting eye-tracking studies offsite, which use integrated webcams to track participants’ gaze in natural settings. This study compares the efficacy of a webcam-based eye tracker to a research-focused screen-mounted eye tracker in code comprehension tasks. We conducted onsite experiments with 49 participants, each using both types of eye trackers simultaneously to assess the webcam-based eye tracker’s capability to capture visual patterns at general, semantic, and token levels and detect individual differences. Additionally, we conducted offsite experiments with 10 participants to supplement the findings. Our findings indicate that while the webcam-based eye tracker effectively captures programmers’ semantic comprehension, but faces challenges in accurately identifying cognitive patterns at a more detailed token level in onsite settings. Furthermore, the elevated noise observed in real-world offsite conditions significantly limits the tracker’s reliability for drawing accurate conclusions. Participants also encountered challenges with calibration and task initiation, highlighting areas for improvement in conducting webcam-based eye-tracking studies offsite in the future.This study investigates the feasibility of webcam-based eye-tracking studies in SE, offers insights to enhance the accuracy of webcam-based eye-tracking in programming potentially, and provides guidelines for future webcam-based eye-tracking study designs. Zihan Fang 0001, Robert Wallace, Zachary Karas, Toby Jia-Jun Li, Collin McMillan, Yu Huang 0015 |
IEEE Trans. Software Eng. | 4 |
| 2025 | HAIPS '25: First ACM CCS Workshop on Human-Centered AI Privacy and SecurityabstractRecent advances in AI/ML create novel and pressing privacy and security challenges—ranging from using generative AI to create harmful content, to the generation of insecure code by AI coding assistants; from oversharing information with ChatGPT to unexpected privacy leaks by LLM agents. At the same time, AI offers new opportunities to address long-standing end-user privacy and security concerns and empower practitioners to adopt better security and privacy practices. In the inaugural workshop of HAIPS'25, we aim to help build and strengthen a community of people enthusiastic about privacy and security issues related to AI from a human-centered perspective, and foster cross-disciplinary research agendas that effectively engage with the human element when addressing these issues. Tianshi Li 0001, Toby Jia-Jun Li, Yaxing Yao, Sauvik Das |
CCS | 2 |
| 2025 | Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive TheoriesabstractUser iterates on their label definitions 5 3 System generates counter examples that vary in single dimensions The system learns pattern rules 2 User labels generated counterexamples 4 Simret Araya Gebreegziabher, Yukun Yang 0008, Elena L. Glassman, Toby Jia-Jun Li |
CHI | 4 |
| 2025 | Hashtag Re-Appropriation for Audience Control on Recommendation-Driven Social Media Xiaohongshu (rednote)abstractAlgorithms have played a central role in personalized recommendations on social media. However, they also present significant obstacles for content creators trying to predict and manage their audience reach. This issue is particularly challenging for marginalized groups seeking to maintain safe spaces. Our study explores how women on Xiaohongshu (rednote), a recommendation-driven social platform, proactively re-appropriate hashtags (e.g., #Baby Supplemental Food) by using them in posts unrelated to their literal meaning. The hashtags were strategically chosen from topics that would be uninteresting to the male audience they wanted to block. Through a mixed-methods approach, we analyzed the practice of hashtag re-appropriation based on 5,800 collected posts and interviewed 24 active users from diverse backgrounds to uncover users' motivations and reactions towards the re-appropriation. This practice highlights how users can reclaim agency over content distribution on recommendation-driven platforms, offering insights into self-governance within algorithmic-centered power structures. Ruyuan Wan, Lingbo Tong, Tiffany Knearem, Toby Jia-Jun Li, Ting-Hao 'Kenneth' Huang, Qunfang Wu |
CHI | 4 |
| 2025 | From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task Automation
Yiwen Yin, Chun Yu, Toby Jia-Jun Li, Aamir Khan Jadoon, Sixiang Cheng, Weinan Shi, Mohan Chen 0007, Yuanchun Shi |
CHI | 4 |
| 2025 | LADICA: A Large Shared Display Interface for Generative AI Cognitive Assistance in Co-located Team CollaborationabstractPeer Reviewed Zheng Zhang 0043, Weirui Peng, Xinyue Chen 0001, Luke Cao, Toby Jia-Jun Li |
CHI | 5 |
| 2025 | CLEAR: Towards Contextual LLM-Empowered Privacy Policy Analysis and Risk Generation for Large Language Model ApplicationsabstractThe rise of end-user applications powered by large language models (LLMs), including both conversational interfaces and add-ons to existing graphical user interfaces (GUIs), introduces new privacy challenges. However, many users remain unaware of the risks. This paper explores methods to increase user awareness of privacy risks associated with LLMs in end-user applications. We conducted five co-design workshops to uncover user privacy concerns and their demand for contextual privacy information within LLMs. Based on these insights, we developed CLEAR (Contextual LLM-Empowered Privacy Policy Analysis and Risk Generation), a just-in-time contextual assistant designed to help users identify sensitive information, summarize relevant privacy policies, and highlight potential risks when sharing information with LLMs. We evaluated the usability and usefulness of CLEAR across two example domains: ChatGPT and the Gemini plugin in Gmail. Our findings demonstrated that CLEAR is easy to use and improves users’ understanding of data practices and privacy risks. We also discussed LLM’s duality in posing and mitigating privacy risks, offering design and policy implications. Daodao Zhou, Yanfang Ye 0001, Toby Jia-Jun Li, Yaxing Yao |
IUI | 4 |
| 2025 | Unequal Opportunities: Examining the Bias in Geographical Recommendations by Large Language Models
Shiran Dudy, Thulasi Tholeti, Resmi Ramachandranpillai, Toby Jia-Jun Li, Ricardo Baeza-Yates |
IUI | 5 |
| 2025 | Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Annalisa Szymanski, Noah Ziems, Heather A. Eicher-Miller, Toby Jia-Jun Li, Meng Jiang 0001, Ronald A. Metoyer |
IUI | 4 |
| 2025 | Careful About What App Promotion Ads Recommend! Detecting and Explaining Malware Promotion via App Promotion Graph
Shang Ma, Shao Yang, Shifu Hou, Toby Jia-Jun Li, Xusheng Xiao, Tao Xie 0001, Yanfang Ye 0001 |
NDSS | 5 |
| 2025 | Why am I seeing this: Democratizing End User Auditing for Online Content Recommendations
Leyang Li, Luke Cao, Yanfang Ye 0001, Tianshi Li 0001, Yaxing Yao, Toby Jia-Jun Li |
UIST | 7 |
| 2025 | AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and VideosabstractVideos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals.In cooking, BLV people often rely on non-visual cues-such as touch, taste, and smell-to navigate their environment, making it difficult to follow UIST '25, September 28-October 01, 2025, Busan, Republic of Korea Ning et al.the predominantly audiovisual instructions found in video recipes.To address this problem, we introduce Aroma, an AI system that provides timely responses to the user based on real-time, contextaware assistance by integrating non-visual cues perceived by the user, a wearable camera feed, and video recipe content.Aroma uses a mixed-initiative approach: it responds to user requests while also proactively monitoring the video stream to offer timely alerts and guidance.This collaborative design leverages the complementary strengths of the user and AI system to align the physical environment with the video recipe, helping the user interpret their current state and make sense of the steps.We evaluated Aroma through a study with eight BLV participants and offered insights for designing interactive AI systems to support BLV individuals in performing ADLs. Zheng Ning, Leyang Li, Daniel Killough, JooYoung Seo, Patrick Carrington, Yapeng Tian, Yuhang Zhao 0001, Franklin Mingzhe Li, Toby Jia-Jun Li |
UIST | 9 |
| 2025 | GLITTER: An AI-assisted Platform for Material-Grounded Asynchronous Discussion in Flipped Learning
Weirui Peng, Yinuo Yang, Zheng Zhang 0043, Toby Jia-Jun Li |
UIST | 4 |
| 2025 | AgentPbD: Interactive Agentic Workflow Generation from User Demonstration on Web BrowsersabstractProgramming by Demonstration (PbD) enables users to automate tasks through examples, but traditional systems generate low-level scripts that are hard to generalize or reuse. Recent advances in Large Language Models (LLMs) offer the potential to infer higher-level task structures, but rely on ambiguous natural language input. We present AgentPbD, a system that synthesizes task-level agentic workflows from a single user demonstration. By capturing browser actions and contextual metadata, AgentPbD automatically infers user goals and intentions, transforming user demonstrations into an editable, modular LLM agent workflow, and displays it on the browser extension interface. Users can further review and modify the workflow through visual programming. We demonstrate how AgentPbD bridges PbD and LLM planning, enabling interpretable and generalizable automation of complex web tasks. Zheng Ning, Toby Jia-Jun Li |
VL/HCC | 4 |
| 2025 | Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code ModificationabstractThis paper presents a study of using large language models (LLMs) in modifying existing code. While LLMs for generating code have been widely studied, their role in code modification remains less understood. Although “prompting” serves as the primary interface for developers to communicate intents to LLMs, constructing effective prompts for code modification introduces challenges different from generation. Prior work suggests that natural language summaries may help scaffold this process, yet such approaches have been validated primarily in narrow domains like SQL rewriting. This study investigates two prompting strategies for LLM-assisted code modification: Direct Instruction Prompting, where developers describe changes explicitly in free-form language, and SummaryMediated Prompting, where changes are made by editing the generated summaries of the code. We conducted an exploratory study with 15 developers who completed modification tasks using both techniques across multiple scenarios. Our findings suggest that developers followed an iterative workflow: understanding the code, localizing the edit, and validating outputs through execution or semantic reasoning. Each prompting strategy presented tradeoffs: direct instruction prompting was more flexible and easier to specify, while summary-mediated prompting supported comprehension, prompt scaffolding, and control. Developers’ choice of strategy was shaped by task goals and context, including urgency, maintainability, learning intent, and code familiarity. These findings highlight the need for more usable prompt interactions, including adjustable summary granularity, reliable summarycode traceability, and consistency in generated summaries. Ningzhi Tang, Emory Smith, Yu Huang 0015, Collin McMillan, Toby Jia-Jun Li |
VL/HCC | 5 |
| 2025 | 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesabstractLarge language models are increasingly applied in real-world scenarios, including research and education. These models, however, come with well-known ethical issues, which may manifest in unexpected ways in human-computer interaction research due to the extensive engagement with human subjects. This paper reports on research practices related to LLM use, drawing on 16 semi-structured interviews and a survey with 50 HCI researchers. We discuss the ways in which LLMs are already being utilized throughout the entire HCI research pipeline, from ideation to system development and paper writing. While researchers described nuanced understandings of ethical issues, they were rarely or only partially able to identify and address those ethical concerns in their own projects. This lack of action and reliance on workarounds was explained through the perceived lack of control and distributed responsibility in the LLM supply chain, the conditional nature of engaging with ethics, and competing priorities. Finally, we reflect on the implications of our findings and present opportunities to shape emerging norms of engaging with large language models in HCI research. Shivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li 0001, Hong Shen 0004 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2025 | Programmer Visual Attention During Context-Aware Code SummarizationabstractProgrammer attention represents the visual focus of programmers on parts of the source code in pursuit of programming tasks. The focus of current research in modeling this programmer attention has been on using mouse cursors, keystrokes, or eye tracking equipment to map areas in a snippet of code. These approaches have traditionally only mapped attention for a single method. However, there is a knowledge gap in the literature because programming tasks such as source code summarization require programmers to use contextual knowledge that can only be found in other parts of the project, not only in a single method. To address this knowledge gap, we conducted an in-depth human study with 10 Java programmers, where each programmer generated summaries for 40 methods from five large Java projects over five one-hour sessions. We used eye tracking equipment to map the visual attention of programmers while they wrote the summaries. We also rate the quality of each summary. We found eye-gaze patterns and metrics that define common behaviors between programmer attention during context-aware code summarization. Specifically, we found that programmers need to read up to 35% fewer words (p$\boldsymbol{ \lt }$0.01) over the whole session, and revisit 13% fewer words (p$ \lt $0.03) as they summarize each method during a session, while maintaining the quality of summaries. We also found that the amount of source code a participant looks at correlates with a higher quality summary, but this trend follows a bell-shaped curve, such that after a threshold reading more source code leads to a significant decrease (p$\boldsymbol{ \lt }$0.01) in the quality of summaries. We also gathered insight into the type of methods in the project that provide the most contextual information for code summarization based on programmer attention. Specifically, we observed that programmers spent a majority of their time looking at methods inside the same class as the target method to be summarized. Surprisingly, we found that programmers spent significantly less time looking at methods in the call graph of the target method. We discuss how our empirical observations may aid future studies towards modeling programmer attention and improving context-aware automatic source code summarization. Robert Wallace, Aakash Bansal, Zachary Karas, Ningzhi Tang, Yu Huang 0015, Toby Jia-Jun Li, Collin McMillan |
IEEE Trans. Software Eng. | 6 |
| 2024 | MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on VideosabstractSpatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized hardware equipment and skills, posing a high barrier for amateur video creators. We present Mimosa, a human-AI co-creation tool that enables amateur users to computationally generate and manipulate spatial audio effects. For a video with only monaural or stereo audio, Mimosa automatically grounds each sound source to the corresponding sounding object in the visual scene and enables users to further validate and fix errors in the location of the sounding objects. Users can also augment the spatial audio effect by flexibly manipulating the sounding source positions and creatively customizing the audio effect. The design of Mimosa exemplifies a human-AI collaboration approach that, instead of utilizing state-of-art end-to-end “black-box” ML models, uses a multistep pipeline that aligns its interpretable intermediate results with the user’s workflow. A lab user study with 15 participants demonstrates Mimosa’s usability, usefulness, expressiveness, and capability in creating immersive spatial audio effects in collaboration with users. Zheng Ning, Zheng Zhang 0043, Jerrick Ban, Ruohong Gan, Yapeng Tian, Toby Jia-Jun Li |
Creativity & Cognition | 7 |
| 2024 | CoCo Matrix: Taxonomy of Cognitive Contributions in Co-writing with Intelligent AgentsabstractIn recent years, there has been a growing interest in employing intelligent agents in writing. Previous work emphasizes the evaluation of the quality of end product—whether it was coherent and polished, overlooking the journey that led to the product, which is an invaluable dimension of the creative process. To understand how to recognize human efforts in co-writing with intelligent writing systems, we adapt Flower and Hayes’ cognitive process theory of writing and propose CoCo Matrix, a two-dimensional taxonomy of entropy and information gain, to depict the new human-agent co-writing model. We define four quadrants and situate thirty-four published systems within the taxonomy. Our research found that low entropy and high information gain systems are under-explored, yet offer promising future directions in writing tasks that benefit from the agent’s divergent planning and the human’s focused translation. CoCo Matrix, not only categorizes different writing systems but also deepens our understanding of the cognitive processes in human-agent co-writing. By analyzing minimal changes in the writing process, CoCo Matrix serves as a proxy for the writer’s mental model, allowing writers to reflect on their contributions. This reflection is facilitated through the measured metrics of information gain and entropy, which provide insights irrespective of the writing system used. Ruyuan Wan, Simret Araya Gebreegziabher, Toby Jia-Jun Li, Karla A. Badillo-Urquiola |
Creativity & Cognition | 3 |
| 2024 | An Empathy-Based Sandbox Approach to Bridge the Privacy Gap among Attitudes, Goals, Knowledge, and BehaviorsabstractManaging privacy to reach privacy goals is challenging, as evidenced by the privacy attitude-behavior gap. Mitigating this discrepancy requires solutions that account for both system opaqueness and users’ hesitations in testing different privacy settings due to fears of unintended data exposure. We introduce an empathy-based approach that allows users to experience how privacy attributes may alter system outcomes in a risk-free sandbox environment from the perspective of artificially generated personas. To generate realistic personas, we introduce a novel pipeline that augments the outputs of large language models (e.g., GPT-4) using few-shot learning, contextualization, and chain of thoughts. Our empirical studies demonstrated the adequate quality of generated personas and highlighted the changes in privacy-related applications (e.g., online advertising) caused by different personas. Furthermore, users demonstrated cognitive and emotional empathy towards the personas when interacting with our sandbox. We offered design implications for downstream applications in improving user privacy literacy. Wenxin Song, Yanfang Ye 0001, Yaxing Yao, Toby Jia-Jun Li |
CHI | 6 |
| 2024 | CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language ModelsabstractCollaborative Qualitative Analysis (CQA) can enhance qualitative analysis rigor and depth by incorporating varied viewpoints. Nevertheless, ensuring a rigorous CQA procedure itself can be both complex and costly. To lower this bar, we take a theoretical perspective to design a one-stop, end-to-end workflow, CollabCoder, that integrates Large Language Models (LLMs) into key inductive CQA stages. In the independent open coding phase, CollabCoder offers AI-generated code suggestions and records decision-making data. During the iterative discussion phase, it promotes mutual understanding by sharing this data within the coding team and using quantitative metrics to identify coding (dis)agreements, aiding in consensus-building. In the codebook development phase, CollabCoder provides primary code group suggestions, lightening the workload of developing a codebook from scratch. A 16-user evaluation confirmed the effectiveness of CollabCoder, demonstrating its advantages over the existing CQA platform. All related materials of CollabCoder, including code and further extensions, will be included in: https://gaojie058.github.io/CollabCoder/. Gionnieve Lim, Tianqin Zhang, Zheng Zhang 0043, Toby Jia-Jun Li, Simon T. Perrault |
CHI | 6 |
| 2024 | SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersabstractBlind or Low-Vision (BLV) users often rely on audio descriptions (AD) to access video content. However, conventional static ADs can leave out detailed information in videos, impose a high mental load, neglect the diverse needs and preferences of BLV users, and lack immersion. To tackle these challenges, we introduce Spica, an AI-powered system that enables BLV users to interactively explore video content. Informed by prior empirical studies on BLV video consumption, Spica offers interactive mechanisms for supporting temporal navigation of frame captions and spatial exploration of objects within key frames. Leveraging an audio-visual machine learning pipeline, Spica augments existing ADs by adding interactivity, spatial sound effects, and individual object descriptions without requiring additional human annotation. Through a user study with 14 BLV participants, we evaluated the usability and usefulness of Spica and explored user behaviors, preferences, and mental models when interacting with augmented ADs. Zheng Ning, Brianna L. Wimer, Keyi Chen 0008, Jerrick Ban, Yapeng Tian, Yuhang Zhao 0001, Toby Jia-Jun Li |
CHI | 8 |
| 2024 | Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationabstractThanks to their generative capabilities, large language models (LLMs) have become an invaluable tool for creative processes. These models have the capacity to produce hundreds and thousands of visual and textual outputs, offering abundant inspiration for creative endeavors. But are we harnessing their full potential? We argue that current interaction paradigms fall short, guiding users towards rapid convergence on a limited set of ideas, rather than empowering them to explore the vast latent design space in generative models. To address this limitation, we propose a framework that facilitates the structured generation of design space in which users can seamlessly explore, evaluate, and synthesize a multitude of responses. We demonstrate the feasibility and usefulness of this framework through the design and development of an interactive system, Luminate, and a user study with 14 professional writers. Our work advances how we interact with LLMs for creative tasks, introducing a way to harness the creative potential of LLMs. Sangho Suh, Meng Chen 0020, Bryan Min, Toby Jia-Jun Li, Haijun Xia |
CHI | 4 |
| 2024 | AutoDroid: LLM-powered Task Automation in AndroidabstractMobile task automation is an attractive technique that aims to enable voice-based hands-free user interaction with smartphones. However, existing approaches suffer from poor scalability due to the limited language understanding ability and the non-trivial manual efforts required from developers or endusers. The recent advance of large language models (LLMs) in language understanding and reasoning inspires us to rethink the problem from a model-centric perspective, where task preparation, comprehension, and execution are handled by a unified language model. In this work, we introduce AutoDroid, a mobile task automation system capable of handling arbitrary tasks on any Android application without manual efforts. The key insight is to combine the commonsense knowledge of LLMs and domain-specific knowledge of apps through automated dynamic analysis. The main components include a functionality-aware UI representation method that bridges the UI with the LLM, exploration-based memory injection techniques that augment the app-specific domain knowledge of LLM, and a multi-granularity query optimization module that reduces the cost of model inference. We integrate AutoDroid with off-the-shelf LLMs including online GPT-4/GPT-3.5 and on-device Vicuna, and evaluate its performance on a new benchmark for memory-augmented Android task automation with 158 common tasks. The results demonstrated that AutoDroid is able to precisely generate actions with an accuracy of 90.9%, and complete tasks with a success rate of 71.3%, outperforming the GPT-4-powered baselines by 36.4% and 39.7%. Hao Wen 0004, Yuanchun Li 0003, Guohong Liu 0002, Shanhui Zhao, Toby Jia-Jun Li, Shiqi Jiang 0002, Yunhao Liu 0001, Yunxin Liu 0001 |
MobiCom | 6 |
| 2024 | SQLucid: Grounding Natural Language Database Queries with Interactive ExplanationsabstractThough recent advances in machine learning have led to significant improvements in natural language interfaces for databases, the accuracy and reliability of these systems remain limited, especially in high-stakes domains. This paper introduces SQLucid, a novel user interface that bridges the gap between non-expert users and complex database querying processes. SQLucid addresses existing limitations by integrating visual correspondence, intermediate query results, and editable step-by-step SQL explanations in natural language to facilitate user understanding and engagement. This unique blend of features empowers users to understand and refine SQL queries easily and precisely. Two user studies and one quantitative experiment were conducted to validate SQLucid’s effectiveness, showing significant improvement in task completion accuracy and user confidence compared to existing interfaces. Our code is available at https://github.com/magic-YuanTian/SQLucid. Jonathan K. Kummerfeld, Toby Jia-Jun Li, Tianyi Zhang 0001 |
UIST | 3 |
| 2024 | Developer Behaviors in Validating and Repairing LLM-Generated Code Using IDE and Eye TrackingabstractThe increasing use of large language model (LLM)-powered code generation tools, such as GitHub Copilot, is transforming software engineering practices. This paper investigates how developers validate and repair code generated by Copilot and examines the impact of code provenance awareness during these processes. We conducted a lab study with 28 participants tasked with validating and repairing Copilot-generated code in three software projects. Participants were randomly divided into two groups: one informed about the provenance of LLM-generated code and the other not. We collected data on IDE interactions, eye-tracking, cognitive workload assessments, and conducted semi-structured interviews. Our results indicate that, without explicit information, developers often fail to identify the LLM origin of the code. Developers exhibit LLM-specific behaviors such as frequent switching between code and comments, different attentional focus, and a tendency to delete and rewrite code. Being aware of the code’s provenance led to improved performance, increased search efforts, more frequent Copilot usage, and higher cognitive workload. These findings enhance our understanding of developer interactions with LLM-generated code and inform the design of tools for effective human-LLM collaboration in software development. Ningzhi Tang, Meng Chen 0020, Zheng Ning, Aakash Bansal, Yu Huang 0015, Collin McMillan, Toby Jia-Jun Li |
VL/HCC | 7 |
| 2024 | Sketchar: Supporting Character Design and Illustration Prototyping Using Generative AIabstractCharacter design in games involves interdisciplinary collaborations, typically between designers who create the narrative content, and illustrators who realize the design vision. However, traditional workflows face challenges in communication due to the differing backgrounds of illustrators and designers, the latter with limited artistic abilities. To overcome these challenges, we created Sketchar, a Generative AI (GenAI) tool that allows designers to prototype game characters and generate images based on conceptual input, providing visual outcomes that can give immediate feedback and enhance communication with illustrators' next step in the design cycle. We conducted a mixed-method study to evaluate the interaction between game designers and Sketchar. We showed that the reference images generated in co-creating with Sketchar fostered refinement of design details and can be incorporated into real-world workflows. Moreover, designers without artistic backgrounds found the Sketchar workflow to be more expressive and worthwhile. This research demonstrates the potential of GenAI in enhancing interdisciplinary collaboration in the game industry, enabling designers to interact beyond their own limited expertise. Long Ling, Ruoyu Wen, Toby Jia-Jun Li, Ray LC |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2024 | From Awareness to Action: Exploring End-User Empowerment Interventions for Dark Patterns in UXabstractThe study of UX dark patterns, i.e., UI designs that seek to manipulate user behaviors, often for the benefit of online services, has drawn significant attention in the CHI and CSCW communities in recent years. To complement previous studies in addressing dark patterns from (1) the designer's perspective on education and advocacy for ethical designs; and (2) the policymaker's perspective on new regulations, we propose an end-user-empowerment intervention approach that helps users (1) raise the awareness of dark patterns and understand their underlying design intents; (2) take actions to counter the effects of dark patterns using a web augmentation approach. Through a two-phase co-design study, including 5 co-design workshops (N=12) and a 2-week technology probe study (N=15), we reported findings on the understanding of users' needs, preferences, and challenges in handling dark patterns and investigated the feedback and reactions to users' awareness of and action on dark patterns being empowered in a realistic in-situ setting. Yuwen Lu, Chao Zhang 0082, Yuewen Yang, Yaxing Yao, Toby Jia-Jun Li |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Insights into Natural Language Database Query Errors: from Attention Misalignment to User Handling StrategiesabstractQuerying structured databases with natural language (NL2SQL) has remained a difficult problem for years. Recently, the advancement of machine learning (ML), natural language processing (NLP), and large language models (LLM) have led to significant improvements in performance, with the best model achieving ∼85% percent accuracy on the benchmark Spider dataset. However, there is a lack of a systematic understanding of the types, causes, and effectiveness of error-handling mechanisms of errors for erroneous queries nowadays. To bridge the gap, a taxonomy of errors made by four representative NL2SQL models was built in this work, along with an in-depth analysis of the errors. Second, the causes of model errors were explored by analyzing the model-human attention alignment to the natural language query. Last, a within-subjects user study with 26 participants was conducted to investigate the effectiveness of three interactive error-handling mechanisms in NL2SQL. Findings from this article shed light on the design of model structure and error discovery and repair strategies for natural language data query interfaces in the future. Zheng Ning, Zheng Zhang 0043, Tianyi Zhang 0001, Toby Jia-Jun Li |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2024 | A Tale of Two Comprehensions? Analyzing Student Programmer Attention during Code SummarizationabstractCode summarization is the task of creating short, natural language descriptions of source code. It is an important part of code comprehension and a powerful method of documentation. Previous work has made progress in identifying where programmers focus in code as they write their own summaries (i.e., Writing). However, there is currently a gap in studying programmers’ attention as they read code with pre-written summaries (i.e., Reading). As a result, it is currently unknown how these two forms of code comprehension compare: Reading and Writing. Also, there is a limited understanding of programmer attention with respect to program semantics. We address these shortcomings with a human eye-tracking study ( n = 27) comparing Reading and Writing. We examined programmers’ attention with respect to fine-grained program semantics, including their attention sequences (i.e., scan paths). We find distinctions in programmer attention across the comprehension tasks, similarities in reading patterns between them, and differences mediated by demographic factors. This can help guide code comprehension in both computer science education and automated code summarization. Furthermore, we mapped programmers’ gaze data onto the Abstract Syntax Tree to explore another representation of human attention. We find that visual behavior on this structure is not always consistent with that on source code. Zachary Karas, Aakash Bansal, Yifan Zhang 0013, Toby Jia-Jun Li, Collin McMillan, Yu Huang 0015 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule SynthesisabstractOver the years, the task of AI-assisted data annotation has seen remarkable advancements. However, a specific type of annotation task, the qualitative coding performed during thematic analysis, has characteristics that make effective human-AI collaboration difficult. Informed by a formative study, we designed PaTAT, a new AI-enabled tool that uses an interactive program synthesis approach to learn flexible and expressive patterns over user-annotated codes in real-time as users annotate data. To accommodate the ambiguous, uncertain, and iterative nature of thematic analysis, the use of user-interpretable patterns allows users to understand and validate what the system has learned, make direct fixes, and easily revise, split, or merge previously annotated codes. This new approach also helps human users to learn data characteristics and form new theories in addition to facilitating the “learning” of the AI model. PaTAT’s usefulness and effectiveness were evaluated in a lab user study. Simret Araya Gebreegziabher, Zheng Zhang 0043, Xiaohang Tang, Yihao Meng, Elena L. Glassman, Toby Jia-Jun Li |
CHI | 6 |
| 2023 | Interactive Text-to-SQL Generation via Editable Step-by-Step ExplanationsabstractRelational databases play an important role in business, science, and more.However, many users cannot fully unleash the analytical power of relational databases, because they are not familiar with database languages such as SQL.Many techniques have been proposed to automatically generate SQL from natural language, but they suffer from two issues: (1) they still make many mistakes, particularly for complex queries, and (2) they do not provide a flexible way for non-expert users to validate and refine incorrect queries.To address these issues, we introduce a new interaction mechanism that allows users to directly edit a stepby-step explanation of a query to fix errors.Our experiments on multiple datasets, as well as a user study with 24 participants, demonstrate that our approach can achieve better performance than multiple SOTA approaches. Zheng Zhang 0043, Zheng Ning, Toby Jia-Jun Li, Jonathan K. Kummerfeld, Tianyi Zhang 0001 |
EMNLP | 4 |
| 2023 | An Empirical Study of Model Errors and User Error Discovery and Repair Strategies in Natural Language Database QueriesabstractRecent advances in machine learning (ML) and natural language processing (NLP) have led to significant improvement in natural language interfaces for structured databases (NL2SQL). Despite the great strides, the overall accuracy of NL2SQL models is still far from being perfect (∼ 75% on the Spider benchmark). In practice, this requires users to discern incorrect SQL queries generated by a model and manually fix them when using NL2SQL models. Currently, there is a lack of comprehensive understanding about the common errors in auto-generated SQLs and the effective strategies to recognize and fix such errors. To bridge the gap, we (1) performed an in-depth analysis of errors made by three state-of-the-art NL2SQL models; (2) distilled a taxonomy of NL2SQL model errors; and (3) conducted a within-subjects user study with 26 participants to investigate the effectiveness of three representative interactive mechanisms for error discovery and repair in NL2SQL. Findings from this paper shed light on the design of future error discovery and repair strategies for natural language data query interfaces. Zheng Ning, Zheng Zhang 0043, Tianyi Zhang 0001, Toby Jia-Jun Li |
IUI | 6 |
| 2023 | Modeling Programmer Attention as Scanpath PredictionabstractThis paper launches a new effort at modeling programmer attention by predicting eye movement scanpaths. Programmer attention refers to what information people intake when performing programming tasks. Models of programmer attention refer to machine prediction of what information is important to people. Models of programmer attention are important because they help researchers build better interfaces, assistive technologies, and more human-like AI. For many years, researchers in SE have built these models based on features such as mouse clicks, key logging, and IDE interactions. Yet the holy grail in this area is scanpath prediction - the prediction of the sequence of eye fixations a person would take over a visual stimulus. A person's eye movements are considered the most concrete evidence that a person is taking in a piece of information. Scanpath prediction is a notoriously difficult problem, but we believe that the emergence of lower-cost, higheraccuracy eye tracking equipment and better large language models of source code brings a solution within grasp. We present an eye tracking experiment with 27 programmers and a prototype scanpath predictor to present preliminary results and obtain early community feedback. Aakash Bansal, Chia-Yi Su, Zachary Karas, Yifan Zhang 0013, Yu Huang 0015, Toby Jia-Jun Li, Collin McMillan |
ASE | 6 |
| 2023 | VISAR: A Human-AI Argumentative Writing Assistant with Visual Programming and Rapid Draft PrototypingabstractIn argumentative writing, writers must brainstorm hierarchical writing goals, ensure the persuasiveness of their arguments, and revise and organize their plans through drafting. Recent advances in large language models (LLMs) have made interactive text generation through a chat interface (e.g., ChatGPT) possible. However, this approach often neglects implicit writing context and user intent, lacks support for user control and autonomy, and provides limited assistance for sensemaking and revising writing plans. To address these challenges, we introduce VISAR, an AI-enabled writing assistant system designed to help writers brainstorm and revise hierarchical goals within their writing context, organize argument structures through synchronized text editing and visual programming, and enhance persuasiveness with argumentation spark recommendations. VISAR allows users to explore, experiment with, and validate their writing plans using automatic draft prototyping. A controlled lab study confirmed the usability and effectiveness of VISAR in facilitating the argumentative writing planning process. Zheng Zhang 0043, Ranjodh Singh Dhaliwal, Toby Jia-Jun Li |
UIST | 4 |
| 2023 | PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual DataabstractAudio-visual learning seeks to enhance the computer’s multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and accessibility, the performance and adoption of existing audio-visual models have been impeded by the availability of high-quality datasets. Annotating audio-visual datasets is laborious, expensive, and time-consuming. To address this challenge, we designed and developed an efficient audio-visual annotation tool called Peanut. Peanut’s human-AI collaborative pipeline separates the multi-modal task into two single-modal tasks, and utilizes state-of-the-art object detection and sound-tagging models to reduce the annotators’ effort to process each frame and the number of manually-annotated frames needed. A within-subject user study with 20 participants found that Peanut can significantly accelerate the audio-visual data annotation process while maintaining high annotation accuracy. Zheng Zhang 0043, Zheng Ning, Chenliang Xu, Yapeng Tian, Toby Jia-Jun Li |
UIST | 5 |
| 2022 | Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionabstractYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Dakuo Wang, Mo Yu, Daniel Ritchie 0002, Bingsheng Yao, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou 0001, Xiaojuan Ma, Diyi Yang, Nanyun Peng 0001, Zhou Yu 0005, Mark Warschauer |
ACL (1) | 8 |
| 2022 | It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story BooksabstractExisting question answering (QA) techniques are created mainly to answer questions asked by humans.But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities.We design an automated question-answer generation (QAG) system for this education scenario: given a story book at the kindergarten to eighth-grade level as input, our system can automatically generate QA pairs that are capable of testing a variety of dimensions of a student's comprehension skills.Our proposed QAG model architecture is demonstrated using a new expert-annotated FairytaleQA dataset, which has 278 child-friendly storybooks with 10,580 QA pairs.Automatic and human evaluations show that our model outperforms stateof-the-art QAG baseline systems.On top of our QAG system, we also start to build an interactive story-telling application for the future real-world deployment in this educational scenario. Bingsheng Yao, Dakuo Wang, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Mo Yu |
ACL (1) | 5 |
| 2022 | StoryBuddy: A Human-AI Collaborative Chatbot for Parent-Child Interactive Storytelling with Flexible Parental InvolvementabstractDespite its benefits for children’s skill development and parent-child bonding, many parents do not often engage in interactive storytelling by having story-related dialogues with their child due to limited availability or challenges in coming up with appropriate questions. While recent advances made AI generation of questions from stories possible, the fully-automated approach excludes parent involvement, disregards educational goals, and underoptimizes for child engagement. Informed by need-finding interviews and participatory design (PD) results, we developed StoryBuddy, an AI-enabled system for parents to create interactive storytelling experiences. StoryBuddy’s design highlighted the need for accommodating dynamic user needs between the desire for parent involvement and parent-child bonding and the goal of minimizing parent intervention when busy. The PD revealed varied assessment and educational goals of parents, which StoryBuddy addressed by supporting configuring question types and tracking child progress. A user study validated StoryBuddy’s usability and suggested design insights for future parent-AI collaboration systems. Zheng Zhang 0043, Bingsheng Yao, Daniel Ritchie 0002, Sherry Tongshuang Wu, Mo Yu, Dakuo Wang, Toby Jia-Jun Li |
CHI | 9 |
| 2021 | Screen2Vec: Semantic Embedding of GUI Screens and GUI ComponentsabstractRepresenting the semantics of GUI screens and components is crucial to data-driven computational methods for modeling user-GUI interactions and mining GUI designs. Existing GUI semantic representations are limited to encoding either the textual content, the visual design and layout patterns, or the app contexts. Many representation techniques also require significant manual data annotation efforts. This paper presents Screen2Vec, a new self-supervised technique for generating representations in embedding vectors of GUI screens and components that encode all of the above GUI features without requiring manual annotation using the context of user interaction traces. Screen2Vec is inspired by the word embedding method Word2Vec, but uses a new two-layer pipeline informed by the structure of GUIs and interaction traces and incorporates screen- and app-specific metadata. Through several sample downstream tasks, we demonstrate Screen2Vec’s key useful properties: representing between-screen similarity through nearest neighbors, composability, and capability to represent user tasks. Toby Jia-Jun Li, Lindsay Popowski, Tom M. Mitchell, Brad A. Myers |
CHI | 1 |
| 2020 | Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsabstractA major problem in task-oriented conversational agents is the lack of support for the repair of conversational breakdowns. Prior studies have shown that current repair strategies for these kinds of errors are often ineffective due to: (1) the lack of transparency about the state of the system's understanding of the user's utterance; and (2) the system's limited capabilities to understand the user's verbal attempts to repair natural language understanding errors. This paper introduces SOVITE, a new multi-modal speech plus direct manipulation interface that helps users discover, identify the causes of, and recover from conversational breakdowns using the resources of existing mobile app GUIs for grounding. SOVITE displays the system's understanding of user intents using GUI screenshots, allows users to refer to third-party apps and their GUI screens in conversations as inputs for intent disambiguation, and enables users to repair breakdowns using direct manipulation on these screenshots. The results from a remote user study with 10 users using SOVITE in 7 scenarios suggested that SOVITE's approach is usable and effective. Toby Jia-Jun Li, Haijun Xia, Tom M. Mitchell, Brad A. Myers |
UIST | 1 |
| 2020 | Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web ApplicationsabstractSupporting voice commands in applications presents significant benefits to users. However, adding such support to existing GUI-based web apps is effort-consuming with a high learning barrier, as shown in our formative study, due to the lack of unified support for creating multi-modal interfaces. We develop Geno---a developer tool for adding the voice input modality to existing web apps without requiring significate NLP expertise. Geno provides a unified workflow for developers to specify functionalities to support by voice (intents), create language models for detecting intents and the relevant information (parameters) from user utterances, and fulfill the intents by either programmatically invoking the corresponding functions or replaying GUI actions on the web app. Geno further supports references to GUI context in voice commands (e.g., "add this to the playlist"). In a study, developers with little NLP expertise were able to add the multi-modal support for two existing web apps using Geno. Ritam Jyoti Sarmah, Yunpeng Ding, Cheuk Yin Phipson Lee, Toby Jia-Jun Li, Xiang 'Anthony' Chen |
UIST | 5 |
| 2020 | Privacy-Preserving Script Sharing in GUI-based Programming-by-Demonstration SystemsabstractAn important concern in end user development (EUD) is accidentally embedding personal information in program artifacts when sharing them. This issue is particularly important in GUI-based programming-by-demonstration (PBD) systems due to the lack of direct developer control of script contents. Prior studies reported that these privacy concerns were the main barrier to script sharing in EUD. We present a new approach that can identify and obfuscate the potential personal information in GUI-based PBD scripts based on the uniqueness of information entries with respect to the corresponding app GUI context. Compared with the prior approaches, ours supports broader types of personal information beyond explicitly pre-specified ones, requires minimal user effort, addresses the threat of re-identification attacks, and can work with third-party apps from any task domain. Our approach also recovers obfuscated fields locally on the script consumer's side to preserve the shared scripts' transparency, readability, robustness, and generalizability. Our evaluation shows that our approach (1) accurately identifies the potential personal information in scripts across different apps in diverse task domains; (2) allows end-user developers to feel comfortable sharing their own scripts; and (3) enables script consumers to understand the operation of shared scripts despite the obfuscated fields. Toby Jia-Jun Li, Brandon Canfield, Brad A. Myers |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2019 | PUMICE: A Multi-Modal Agent that Learns Concepts and Conditionals from Natural Language and DemonstrationsabstractNatural language programming is a promising approach to enable end users to instruct new tasks for intelligent agents. However, our formative study found that end users would often use unclear, ambiguous or vague concepts when naturally instructing tasks in natural language, especially when specifying conditionals. Existing systems have limited support for letting the user teach agents new concepts or explaining unclear concepts. In this paper, we describe a new multi-modal domain-independent approach that combines natural language programming and programming-by-demonstration to allow users to first naturally describe tasks and associated conditions at a high level, and then collaborate with the agent to recursively resolve any ambiguities or vagueness through conversations and demonstrations. Users can also define new procedures and concepts by demonstrating and referring to contents within GUIs of existing mobile apps. We demonstrate this approach in PUMICE, an end-user programmable agent that implements this approach. A lab study with 10 users showed its usability. Toby Jia-Jun Li, Marissa Radensky, Justin Jia, Kirielle Singarajah, Tom M. Mitchell, Brad A. Myers |
UIST | 1 |
| 2018 | Kite: Building Conversational Bots from Mobile AppsabstractTask-oriented chatbots allow users to carry out tasks (e.g., ordering a pizza) using natural language conversation. The widely-used slot-filling approach for building bots of this type requires significant hand-coding, which hinders scalability. Recently, neural network models have been shown to be capable of generating natural "chitchat" conversations, but it is unclear whether they will ever work for task modeling. Kite is a practical system for bootstrapping task-oriented bots, leveraging both approaches above. Kite's key insight is that while bots encapsulate the logic of user tasks into conversational forms, existing apps encapsulate the logic of user tasks into graphical user interfaces. A developer demonstrates a task using a relevant app, and from the collected interaction traces Kite automatically derives a task model, a graph of actions and associated inputs representing possible task execution paths. A task model represents the logical backbone of a bot, on which Kite layers a question-answer interface generated using a hybrid rule-based and neural network approach. Using Kite, developers can automatically generate bot templates for many different tasks. In our evaluation, it extracted accurate task models from 25 popular Android apps spanning 15 tasks. Appropriate questions and high-quality answers were also generated. Our developer study suggests that developers, even without any bot developing experience, can successfully generate bot templates using Kite. Toby Jia-Jun Li, Oriana Riva |
MobiSys | 1 |
| 2018 | APPINITE: A Multi-Modal Interface for Specifying Data Descriptions in Programming by Demonstration Using Natural Language InstructionsabstractA key challenge for generalizing programming-by-demonstration (PBD) scripts is the data description problem - when a user demonstrates performing an action, the system needs to determine features for describing this action and the target object in a way that can reflect the user's intention for the action. However, prior approaches for creating data descriptions in PBD systems have problems with usability, applicability, feasibility, transparency and/or user control. Our APPINITE system introduces a multimodal interface with which users can specify data descriptions verbally using natural language instructions. APPINITE guides users to describe their intentions for the demonstrated actions through mixed-initiative conversations. APPINITE constructs data descriptions for these actions from the natural language instructions. Our evaluation showed that APPINITE is easy-to-use and effective in creating scripts for tasks that would otherwise be difficult to create with prior PBD systems, due to ambiguous data descriptions in demonstrations on GUIs. Toby Jia-Jun Li, Igor Labutov, Xiaohan Nancy Li, Xiaoyi Zhang 0006, Wenze Shi, Wanling Ding, Tom M. Mitchell, Brad A. Myers |
VL/HCC | 1 |
| 2018 | How End Users Express Conditionals in Programming by Demonstration for Mobile AppsabstractThough conditionals are an integral component of programming, providing an easy means of creating conditionals remains a challenge for programming-by-demonstration (PBD) systems for task automation. We hypothesize that a promising method for implementing conditionals in such systems is to incorporate the use of verbal instructions. Verbal instructions supplied concurrently with demonstrations have been shown to improve the generalizability of PBD. However, the challenge of supporting conditional creation using this multi-modal approach has not been addressed. In this extended abstract, we present our study on understanding how end users describe conditionals in natural language for mobile app tasks. We conducted a formative study of 56 participants asking them to verbally describe conditionals in different settings for 9 sample tasks and to invent conditional tasks. Participant responses were analyzed using open coding and revealed that, in the context of mobile apps, end users often omit desired else statements when explaining conditionals, sometimes use ambiguous concepts in expressing conditionals, and often desire to implement complex conditionals. Based on these findings, we discuss the implications for designing a multimodal PBD interface to support the creation of conditionals. Marissa Radensky, Toby Jia-Jun Li, Brad A. Myers |
VL/HCC | 2 |
| 2017 | SUGILITE: Creating Multimodal Smartphone Automation by DemonstrationabstractSUGILITE is a new programming-by-demonstration (PBD) system that enables users to create automation on smartphones. SUGILITE uses Android's accessibility API to support automating arbitrary tasks in any Android app (or even across multiple apps). When the user gives verbal commands that SUGILITE does not know how to execute, the user can demonstrate by directly manipulating the regular apps' user interface. By leveraging the verbal instructions, the demonstrated procedures, and the apps? UI hierarchy structures, SUGILITE can automatically generalize the script from the recorded actions, so SUGILITE learns how to perform tasks with different variations and parameters from a single demonstration. Extensive error handling and context checking support forking the script when new situations are encountered, and provide robustness if the apps change their user interface. Our lab study suggests that users with little or no programming knowledge can successfully automate smartphone tasks using SUGILITE. Toby Jia-Jun Li, Amos Azaria, Brad A. Myers |
CHI | 1 |
| 2017 | End user mobile task automation using multimodal programming by demonstrationabstractConversational agents are often used to perform tasks on smartphones, but existing conversational agents are limited in capabilities and lack of customizability. My work explores using the programming-by-demonstration approach to enable end users to program new tasks for conversational agents by demonstrating using the familiar graphical user interfaces of third-party apps. I propose to use a multi-modal (demonstration and verbal instruction) interface to support generalization, editing, error handling as well as creating control structures in creating such smartphone automation. Toby Jia-Jun Li |
VL/HCC | 1 |
| 2016 | Not at Home on the Range: Peer Production and the Urban/Rural DivideabstractWikipedia articles about places, OpenStreetMap features, and other forms of peer-produced content have become critical sources of geographic knowledge for humans and intelligent technologies. In this paper, we explore the effectiveness of the peer production model across the rural/urban divide, a divide that has been shown to be an important factor in many online social systems. We find that in both Wikipedia and OpenStreetMap, peer-produced content about rural areas is of systematically lower quality, is less likely to have been produced by contributors who focus on the local area, and is more likely to have been generated by automated software agents (i.e. "bots"). We then codify the systemic challenges inherent to characterizing rural phenomena through peer production and discuss potential solutions. Isaac L. Johnson, Allen Yilun Lin, Toby Jia-Jun Li, Andrew Hall, Aaron Halfaker, Johannes Schöning, Brent J. Hecht |
CHI | 3 |
| 2014 | Leveraging advances in natural language processing to better understand Tobler's first law of geographyabstractTobler's First Law of Geography (TFL) is one of the key reasons why "spatial is special". The law, which states that "everything is related to everything else, but near things are more related than distant things", is central to the management, presentation, and analysis of geographic information. However, despite the importance of TFL, we have a limited general understanding of its domain-neutral properties. In this paper, we leverage recent advances in the natural language processing domain of semantic relatedness estimation to, for the first time, robustly evaluate the extent to which relatedness between spatial entities decreases over distance in a domain-neutral fashion. Our results reveal that, in general, TFL can indeed be considered a globally recognized domain-neutral property of geographic information but that there is a distance beyond which being nearer, on average, no longer means being more related. Toby Jia-Jun Li, Shilad Sen, Brent J. Hecht |
SIGSPATIAL/GIS | 1 |
| 2014 | WikiBrain: Democratizing computation on WikipediaabstractWikipedia is known for serving humans' informational needs. Over the past decade, the encyclopedic knowledge encoded in Wikipedia has also powerfully served computer systems. Leading algorithms in artificial intelligence, natural language processing, data mining, geographic information science, and many other fields analyze the text and structure of articles to build computational models of the world. Shilad Sen, Toby Jia-Jun Li, Brent J. Hecht |
OpenSym | 2 |