Xingyu Liu 0002

dblp:88/10181-2 · also Xingyu Bruce Liu · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-6988-5471ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 11 · 5 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
abstract
LLMs are now embedded in a wide range of everyday scenarios. However, their inherent hallucinations risk hiding misinformation in fluent responses, raising concerns about overreliance on AI. Detecting overreliance is challenging, as it often arises in complex, dynamic contexts and cannot be easily captured by post-hoc task outcomes. In this work, we aim to investigate how users’ behavioral patterns correlate with overreliance. We collected interaction logs from 77 participants working with an LLM injected plausible misinformation across three real-world tasks and we assessed overreliance by whether participants detected and corrected these errors. By semantically encoding and clustering segments of user interactions, we identified five behavioral patterns linked to overreliance: users with low overreliance show careful task comprehension and fine-grained navigation; users with high overreliance show frequent copy-paste, skipping initial comprehension, repeated LLM references, coarse locating, and accepting misinformation despite hesitation. We discuss design implications for mitigation.
Chang Liu 0150, Qinyi Zhou, Xinjie Shen, Xingyu Liu 0002, Sherry Tongshuang Wu, Xiang 'Anthony' Chen
CHI4
2026 Unraveling multiparty conversations: From human interaction mechanisms to conversational agent challenges and persona design
abstract
Multiparty conversations are ubiquitous and indispensable in diverse social and collaborative contexts. However, current conversational agents (CAs) face significant challenges in effectively engaging in such interactions, particularly within text-based environments. While earlier limitations were often attributed to the inadequacies of AI models, recent advances in large language models now compel us to revisit both our understanding of multiparty conversation and the way we design CAs. This paper synthesizes findings from two complementary qualitative investigations and proposes a conceptual model for designing CAs that can genuinely participate, rather than merely function as tools or outsiders. The first study, employing retrospective think-aloud sessions (N=30) with users in text-based multiparty settings, uncovers 5 key interactional mechanisms (e.g., Turn-taking Management, Presence Management) that underpin successful human-human multiparty interactions, derived from participants’ articulated perceptions and reasoning. Subsequently, the second study, through semi-structured interviews (N=15), identifies user expectations for CA integration and key traits (e.g., proactivity, social authenticity) that shape an ideal CA persona perceived by users as a genuine participant. Drawing from these human-centric insights, we then derive design considerations, aiming to guide the development of CAs capable of more natural, effective, and socially intelligent participation in multiparty conversation.
Shitao Fang, Xingyu Liu 0002, Takeo Igarashi, Koji Yatani
Int. J. Hum. Comput. Stud.2
2025 Proactive Conversational Agents with Inner Thoughts
Xingyu Liu 0002, Shitao Fang, Weiyan Shi 0001, Chien-Sheng Wu, Takeo Igarashi, Xiang 'Anthony' Chen
CHI1
2025 CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive Commenting
abstract
Figure 1: CoSight, a Chrome extension developed as a design probe to explore how lightweight interface nudges might encourage accessibility contributions from sighted video viewers when watching and commenting.The prototype augments YouTube video pages with features inspired by Fogg's Behavior Model [24], including color labels to highlight accessibility gaps (sparks), hints and references to guide contributions (facilitators), and reminders at key moments (signals).
Ruolin Wang, Xingyu Liu 0002, Wayne Zhang 0004, Ziqian Liao, Ziwen Li 0001, Amy Pavel, Xiang 'Anthony' Chen
UIST2
2024 Human I/O: Towards a Unified Approach to Detecting Situational Impairments
abstract
Situationally Induced Impairments and Disabilities (SIIDs) can significantly hinder user experience in contexts such as poor lighting, noise, and multi-tasking. While prior research has introduced algorithms and systems to address these impairments, they predominantly cater to specific tasks or environments and fail to accommodate the diverse and dynamic nature of SIIDs. We introduce Human I/O, a unified approach to detecting a wide range of SIIDs by gauging the availability of human input/output channels. Leveraging egocentric vision, multimodal sensing and reasoning with large language models, Human I/O achieves a 0.22 mean absolute error and a 82% accuracy in availability prediction across 60 in-the-wild egocentric video recordings in 32 different scenarios. Furthermore, while the core focus of our work is on the detection of SIIDs rather than the creation of adaptive user interfaces, we showcase the efficacy of our prototype via a user study with 10 participants. Findings suggest that Human I/O significantly reduces effort and improves user experience in the presence of SIIDs, paving the way for more adaptive and accessible interactive systems in the future.
Xingyu Liu 0002, Jiahao Nick Li, David Kim 0002, Xiang 'Anthony' Chen, Ruofei Du
CHI1
2023 Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual Programming
abstract
In recent years, there has been a proliferation of multimedia applications that leverage machine learning (ML) for interactive experiences. Prototyping ML-based applications is, however, still challenging, given complex workflows that are not ideal for design and experimentation. To better understand these challenges, we conducted a formative study with seven ML practitioners to gather insights about common ML evaluation workflows.
Ruofei Du, Na Li 0034, Michelle Carney, Scott Miles, Maria Kleiner, Xiuxiu Yuan, Yinda Zhang 0001, Anuva Kulkarni, Xingyu Liu 0002, Ahmed Sabie, Sergio Orts, Abhishek Kar, Ram Iyengar, Adarsh Kowdle, Alex Olwal
CHI10
2023 Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals
abstract
Video conferencing solutions like Zoom, Google Meet, and Microsoft Teams are becoming increasingly popular for facilitating conversations, and recent advancements such as live captioning help people better understand each other. We believe that the addition of visuals based on the context of conversations could further improve comprehension of complex or unfamiliar concepts. To explore the potential of such capabilities, we conducted a formative study through remote interviews (N=10) and crowdsourced a dataset of over 1500 sentence-visual pairs across a wide range of contexts. These insights informed Visual Captions, a real-time system that integrates with a video conferencing platform to enrich verbal communication. Visual Captions leverages a fine-tuned large language model to proactively suggest relevant visuals in open-vocabulary conversations. We present findings from a lab study (N=26) and an in-the-wild case study (N=10), demonstrating how Visual Captions can help improve communication through visual augmentation in various scenarios.
Xingyu Liu 0002, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang 'Anthony' Chen, Ruofei Du
CHI1
2023 Social Wormholes: Exploring Preferences and Opportunities for Distributed and Physically-Grounded Social Connections
abstract
Ubiquitous computing encapsulates the idea for technology to be interwoven into the fabric of everyday life. As computing blends into everyday physical artifacts, powerful opportunities open up for social connection. Prior connected media objects span a broad spectrum of design combinations. Such diversity suggests that people have varying needs and preferences for staying connected to one another. However, since these designs have largely been studied in isolation, we do not have a holistic understanding around how people would configure and behave within a ubiquitous social ecosystem of physically-grounded artifacts. In this paper, we create a technology probe called Social Wormholes, that lets people configure their own home ecosystem of connected artifacts. Through a field study with 24 participants, we report on patterns of behaviors that emerged naturally in the context of their daily lives and shine a light on how ubiquitous computing could be leveraged for social computing.
Joanne Leong, Yuanyang Teng, Xingyu Liu 0002, Hanseul Jun, Sven Kratz, Yu Jiang Tham, Andrés Monroy-Hernández, Brian A. Smith 0001, Rajan Vaish
Proc. ACM Hum. Comput. Interact.3
2022 CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding
abstract
Authors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers and captioners, due to the difficulty of identifying accessibility problems in videos. A video author will have to watch the video through and manually check for inaccessible information frame-by-frame, for both visual and auditory modalities. In this paper, we present CrossA11y, a system that helps authors efficiently detect and address visual and auditory accessibility issues in videos. Using cross-modal grounding analysis, CrossA11y automatically measures accessibility of visual and audio segments in a video by checking for modality asymmetries. CrossA11y then displays these segments and surfaces visual and audio accessibility issues in a unified interface, making it intuitive to locate, review, script AD/CC in-place, and preview the described and captioned video immediately. We demonstrate the effectiveness of CrossA11y through a lab study with 11 participants, comparing to existing baseline.
Xingyu Liu 0002, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel
UIST1
2021 What Makes Videos Accessible to Blind and Visually Impaired People?
abstract
User-generated videos are an increasingly important source of information online, yet most online videos are inaccessible to blind and visually impaired (BVI) people. To find videos that are accessible, or understandable without additional description of the visual content, BVI people in our formative studies reported that they used a time-consuming trial-and-error approach: clicking on a video, watching a portion, leaving the video, and repeating the process. BVI people also reported video accessibility heuristics that characterize accessible and inaccessible videos. We instantiate 7 of the identified heuristics (2 audio-related, 2 video-related, and 3 audio-visual) as automated metrics to assess video accessibility. We collected a dataset of accessibility ratings of videos by BVI people and found that our automatic video accessibility metrics correlated with the accessibility ratings (Adjusted R2 = 0.642). We augmented a video search interface with our video accessibility metrics and predictions. BVI people using our augmented video search interface selected an accessible video more efficiently than when using the original search interface. By integrating video accessibility metrics, video hosting platforms could help people surface accessible videos and encourage content creators to author more accessible products, improving video accessibility for all.
Xingyu Liu 0002, Patrick Carrington, Xiang 'Anthony' Chen, Amy Pavel
CHI1
2019 Making Memes Accessible
abstract
Images on social media platforms are inaccessible to people with vision impairments due to a lack of descriptions that can be read by screen readers. Providing accurate alternative text for all visual content on social media is not yet feasible, but certain subsets of images, such as internet memes, offer affordances for automatic or semi-automatic generation of alternative text. We present two methods for making memes accessible semi-automatically through (1) the generation of rich alternative text descriptions and (2) the creation of audio macro memes. Meme authors create alternative text templates or audio meme templates, and insert placeholders instead of the meme text. When a meme with the same image is encountered again, it is automatically recognized from a database of meme templates. Text is then extracted and either inserted into the alternative text template or rendered in the audio template using text-to-speech. In our evaluation of meme formats with 10 Twitter users with vision impairments, we found that most users preferred alternative text memes because the description of the visual content conveys the emotional tone of the character. As the preexisting templates can be automatically matched to memes using the same visual image, this combined approach can make a large subset of images on the web accessible, while preserving the emotion and tone inherent in the image memes.
Cole Gleason, Amy Pavel, Xingyu Liu 0002, Patrick Carrington, Lydia B. Chilton, Jeffrey P. Bigham
ASSETS3