Wei Xiang 0008

dblp:37/1682-8 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-2058-5379ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 14 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round Debate
abstract
Automating Graphical User Interface (GUI) operations with Multimodal Large Language Models (MLLMs) is promising but remains bottlenecked in real-world long-horizon settings. Key challenges include ensuring precise grounding across diverse interfaces and handling irreversible errors in extended workflows. Current methods often struggle to distinguish targets in low Signal-to-Noise Ratio (SNR) environments and lack sufficient pre-execution verification to prevent error accumulation. To address this, we propose the Memory-augmented Debate System (MaDS). Specifically, MaDS combines: (1) a Dual-Layer Memory Module that integrates universal interaction priors with scenario-specific operational experience to mitigate grounding hallucinations; and (2) Multi-Round Debate that performs pre-execution verification, while transforming execution failures into retrievable Negative Warnings to reduce repeated errors. Additionally, we introduce MaDS-Benchmark, a benchmark for long-horizon mobile GUI tasks with process-oriented evaluation. Experiments show that MaDS achieves a 90.23% Task Success Rate on MaDS-Benchmark and strong performance on public benchmarks including AITW, AITZ, CAGUI, and GUIOdyssey.
Pengchen Chen, Shi Chen 0005, Qiming Ye, Xinli Chen, Wei Xiang 0008
ACL (1)6
2026 Pika: Designing a social-support agent to improve drivers' experience in gig work
Wei Xiang 0008, Xinli Chen, Tianhui Guo, Shi Chen 0005
Int. J. Hum. Comput. Stud.1
2026 Visionary Co-Driver: Enhancing Driver Perception of Potential Risks With LLM and HUD
abstract
Drivers’ perception of risky situations has always been a challenge in driving. Existing risk-detection methods excel at identifying collisions but face challenges in assessing the behavior of road users in non-collision situations. This paper introduces Visionary Co-Driver, a system that leverages large language models (LLMs) to identify non-collision roadside risks and alert drivers based on their eye movements. Specifically, the system combines video processing algorithms and LLMs to identify potentially risky road users. These risks are dynamically indicated on an adaptive heads-up display interface to enhance drivers’ attention. A user study with 41 drivers confirms that Visionary Co-Driver improves drivers’ risk perception and supports their recognition of roadside risks.
Wei Xiang 0008, Ziyue Lei, Lingyun Sun
IEEE Trans. Intell. Transp. Syst.1
2025 Voice by the Non-sighted: Practices and Challenges of Audiobook Voice Actors with Blind and Low Vision in China
Shi Chen 0005, Jingao Zhang, Suqi Lou, Wei Xiang 0008, Lingyun Sun
CHI5
2025 Hand by Hand: LLM Driving EMS Assistant for Operational Skill Learning
abstract
Operational skill learning, inherently physical and reliant on hands-on practice and kinesthetic feedback, has yet to be effectively replicated in large language model (LLM)-supported training. Current LLM training assistants primarily generate customized textual feedback, neglecting the crucial kinesthetic modality. This gap derives from the textual and uncertain nature of LLMs, compounded by concerns on user acceptance of LLM driven body control. To bridge this gap and realize the potential of collaborative human-LLM action, this work explores human experience of LLM driven kinesthetic assistance. Specifically, we introduced an "Align-Analyze-Adjust" strategy and developed FlightAxis, a tool that integrates LLM with Electrical Muscle Stimulation (EMS) for flight skill acquisition, a representative operational skill domain. FlightAxis learns flight skills from manuals and guides forearm movements during simulated flight tasks. Our results demonstrate high user acceptance of LLM-mediated body control and significantly reduced task completion times. Crucially, trainees reported that this kinesthetic assistance enhanced their awareness of operation flaws and fostered increased engagement in the training process, rather than relieving perceived load. This work demonstrated the potential of kinesthetic LLM training in operational skill acquisition.
Wei Xiang 0008, Ziyue Lei, Haoyuan Che, Fangyuan Ye, Xueting Wu, Lingyun Sun
IJCAI1
2025 CareEmo: Supporting Caregivers with Personalized Communication Approaches to Enhance Older Adults' Emotional Well-being
abstract
Memory plays a crucial role in caregiving, helping caregivers to address the emotional needs of older adult care recipients. However, this is the burden, especially when caregivers need to serve multiple care recipients. This study presents CareEmo, an assistant that delivers memory and emotional support to caregivers via Bluetooth earbuds, enabling personalized communication and enhancing care recipients’ emotional well-being. Through a formative study that included field observations and interviews with caregivers, family members and care recipients, we designed modules: emotion recognition, user profile creation, and care advice provision. CareEmo identifies emotional changes in care recipients, summarizes their emotional needs and interests, then provides caregivers with appropriate advice based on recipients’ memories and caregiving histories. An empirical study involving 16 participants (8 caregivers and 8 care recipients), showed that CareEmo improved both caregiver and care recipient emotional experience and contributed a high quality of emotional caregiving.
Muchen Li, Wei Xiang 0008, Mengyun Jiang, Xueting Wu
SMC2
2025 Driver Assistant: Persuading Drivers to Adjust Secondary Tasks Using Large Language Models
abstract
Level 3 automated driving systems allows drivers to engage in secondary tasks while diminishing their perception of risk. In the event of an emergency necessitating driver intervention, the system will alert the driver with a limited window for reaction and imposing a substantial cognitive burden. To address this challenge, this study employs a Large Language Model (LLM) to assist drivers in maintaining an appropriate attention on road conditions through a "humanized" persuasive advice. Our tool leverages the road conditions encountered by Level 3 systems as triggers, proactively steering driver behavior via both visual and auditory routes. Empirical study indicates that our tool is effective in sustaining driver attention with reduced cognitive load and coordinating secondary tasks with takeover behavior. Our work provides insights into the potential of using LLMs to support drivers during multi-task automated driving.
Wei Xiang 0008, Muchen Li, Manling Zheng, Hanfei Zhu, Mengyun Jiang, Lingyun Sun
SMC1
2025 SocializeChat: A GPT-Based AAC Tool Grounded in Personal Memories to Support Social Communication
abstract
Elderly people with speech impairments often face challenges in engaging in meaningful social communication, particularly when using Augmentative and Alternative Communication (AAC) tools that primarily address basic needs. Moreover, effective chats often rely on personal memories, which is hard to extract and reuse. We introduce SocializeChat, an AAC tool that generates sentence suggestions by drawing on users’ personal memory records. By incorporating topic preference and interpersonal closeness, the system reuses past experience and tailors suggestions to different social contexts and conversation partners. SocializeChat not only leverages past experiences to support interaction, but also treats conversations as opportunities to create new memories, fostering a dynamic cycle between memory and communication. A user study shows its potential to enhance the inclusivity and relevance of AAC-supported social interaction.
Wei Xiang 0008, Yunkai Xu, Yuyang Fang, Zhuyu Teng, Zhaoqu Jiang, Beijia Hu, Jinguo Yang
SMC1
2025 VAEnvGen: A Real-Time Virtual Agent Environment Generation System Based on Large Language Models
abstract
Environment plays an important role in non-verbal communication for human-virtual agent interaction. Existing research explores the influence of an agent’s appearance and attributes to enhance human-virtual agent communication. However, there is no common practice for dynamically adjusting the surrounding environments of the virtual agent. In this paper, we introduce a real-time virtual agent environment generation system (VAEnvGen), which contributes to the field by enhancing users’ content perception and improving task performance through dynamic environment adjustment. The system dynamically analyzes both the appropriate communication environment and filters the key information according to the current context. Leveraging Large Language Models, it generates a pseudo-3D background space to create an engaging atmosphere and a dynamic foreground content space for vivid key information display, thereby significantly enhancing content perception. For widespread adoption and flexibility, VAEnvGen is developed as a web application. We further evaluate the impact of VAEnvGen on content perception, user attention, and subjective satisfaction through a mixed-design user study with 50 participants. Quantitative and qualitative results reveal significant improvements in content perception, task completion time, and user satisfaction when using VAEnvGen. The system effectively redistributes user attention from subtitles and the virtual agent itself to the dynamically generated background and key foreground information, leading to a more immersive and less fatiguing user experience.
Jingyu Wu, Pengchen Chen, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun
Int. J. Hum. Comput. Interact.4
2024 SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications
abstract
The conflict between the rapid iteration demand of prototyping and the time-consuming nature of user tests has led researchers to adopt AI methods to identify usability issues. However, these AI-driven methods concentrate on evaluating the feasibility of a system, while often overlooking the influence of specified user characteristics and usage contexts. Our work proposes a tool named SimUser based on large language models (LLMs) with the Chain-of-Thought structure and user modeling method. It generates usability feedback by simulating the interaction between users and applications, which is influenced by user characteristics and contextual factors. The empirical study (48 human users and 21 designers) validated that in the context of a simple smartwatch interface, SimUser could generate heuristic usability feedback with the similarity varying from 35.7% to 100% according to the user groups and usability category. Our work provides insights into simulating users by LLM to improve future design activities.
Wei Xiang 0008, Hanfei Zhu, Suqi Lou, Xinli Chen, Zhenghua Pan, Yuping Jin, Shi Chen 0005, Lingyun Sun
CHI1
2024 CNAMD Corpus: A Chinese Natural Audiovisual Multimodal Database of Conversations for Social Interactive Agents
abstract
Impressive progress has been made in developing companion Socially Interactive Agents (SIAs) that provide companionship and reduce loneliness. However, recent works focus on analyzing multimodal feedback in Answer part but ignore Question part. Furthermore, research on SIAs is primarily based on English, which poses a challenge for Chinese SIAs because of cultural differences between English and Chinese. Therefore, we introduce a Chinese Natural Audiovisual Multimodal Database (CNAMD) corpus, the first and largest freely available Chinese multimodal database for multi-person interaction, containing 48 hours of videos and annotations across eight modalities. Using CNAMD, we analyze the characteristics of vocal-verbal, audio, behavioral, and multimodal combinations during questioning, test the performance of six baselines on three tasks, and propose improvements for processing daily Chinese data. The present findings will help designers consider Chinese customs and language when designing Chinese SIAs, making them more suitable for the Chinese cultural context and users.
Jingyu Wu, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun, Hongzeng Zhang, Yanxu Li
Int. J. Hum. Comput. Interact.3
2024 Differed Risk Perception in Manual and Automated Driving: An Empirical Study of Varied Conditions
abstract
The interface of L3 automated driving systems needs to take both automated and manual driving into consideration. This calls for a comprehension of drivers’ differed risk perception under manual and automated driving conditions, thus supporting description of contextual information and appropriate risk warnings. Existing studies have reported drivers’ impaired ability during automated driving, however, a quantified measurement is still lacking. This study tried to measure the difference in drivers’ ability to perceive risks during automated and manual driving. Specifically, a simulated driving experiment in car-following scenarios was conducted to collect drivers’ perceived risk under multiple manual and automated driving conditions, including varied motion directions, speed, and distance among vehicles. Then, the influences of driving mode, motion directions, speed, and distance on drivers’ risk perceptions were described using a linear mixed model. The result demonstrated a complicated interaction effect. Automated driving impaired drivers’ risk perception, and this effect was less severe in highly risky events. In both manual and automated driving, drivers were less sensitive to risk only when risky events happened backwards. These results indicated drivers’ varied ability under multiple conditions, and supported warning design and interface refinement under automated and manual driving conditions.
Wei Xiang 0008
Int. J. Hum. Comput. Interact.1
2023 Dangerous Slime: A Game for Improving Situation Awareness in Automated Driving
abstract
The role of driver is changing from controller to regulator as automated vehicles become more common. This change leads to deceased situation awareness (SA) because of passive engagement and increased non-driving related tasks (NDRT), and might affected drivers' takeover. This study designed a gamified prototype named Dangerous Slime to help drivers maintain SA during automated driving. Dangerous Slime turned surrounding cars into slimes that attacked drivers' cars; drivers need to response to these attacks, which increased their attention to nearby objects. Drivers can take the game as non-driving related tasks (NDRTs) during the whole automated driving process. When compared to NDRT of watching films, the game achieved an improvement in drivers' SA and positive user feedback. Moreover, the styles of games affect drivers' behavior and SA. This study revealed how the game influenced drivers' SA, and took a step toward improving the safety of automated driving in a pleasant way.
Linhao Ye, Xianzhe Zheng, Hanfei Zhu, Wei Xiang 0008
AutomotiveUI5
2023 Supporting Crowd Workers in Ideation Tasks Through Information Gathering and Reflective Activity
abstract
Crowdsourcing is widely used to solve creative problems of ideation in HCI domain. To improve the creativity of crowdsourcing outcomes, researchers have proposed multiple approaches to support crowd workers’ idea proposal step. Other steps in designers’ creative process, such as information gathering and reflective activity, also impact idea creativity, while their effects in crowdsourcing scenarios remain unexplored. Therefore, referring to the creativity research, this study proposed an approach that optimized crowdsourcing tasks to help crowd workers perform information gathering and reflective activity. Two experiments involving 427 workers were conducted to test the effects of our approach. Results showed that instructing crowd workers to gather information and reflect on their ideas positively affected the creativity of crowdsourcing outcomes. This study provides inspiration for the optimization of crowdsourcing tasks and offers insights to the researchers who focus on using crowd power to achieve the ideation process.
Lingyun Sun, Wei-yue Gao, Wei Xiang 0008
Int. J. Hum. Comput. Interact.3
2022 Enrichment of Product Presentation Video: Methods and Impacts on User Experience
abstract
Product presentation video (PPV) is a genre of short-form video in online retailing that presents product features and facilitates online shopping. To improve PPVs, quality and provide a better experience for viewers, PPV producers, most of whom are online retailers and non-professionals in video production, have made efforts to enrich these short-form videos. Despite the great demand for PPV enrichment, these methods have not been systematically explored, and their impacts on user experience remain unclear, impeding the improvement of PPVs. This study combined qualitative and quantitative methods to explore the impacts of PPV enrichment methods on user experience. As an exploratory study, we focused on PPVs of female fashion products on Chinese e-commerce platforms. We collected 240 PPVs and summarized ten enrichment methods from them accordingly. A questionnaire-based experiment, including 48 participants, was then conducted to explore the impacts of these methods on the experience of PPVs. Results indicated that eight out of ten methods effectively improved PPV experience from multiple dimensions. This study brings insights for exploring PPV enrichment from the perspective of user experience and provides support for PPV production process.
Wei-yue Gao, Wei Xiang 0008, Xuanhui Liu, Lingyun Sun
QoMEX2
2022 Impacts of Presenting Extra Information in Short Videos via Text and Voice on User Experience
abstract
Short video is an increasingly prevalent medium in online shopping environments to present products. To cope with the great demand for short videos rising from the enormous number and the rapid update of online products, computer-supported video production is becoming a trend. The optimization of short videos considering user experience is essential. Currently, using text and voice to integrate extra information into short videos is a potential and promising approach for optimizing computer-supported video production, while the effects of these elements on user experience remain unclear. In this study, we conducted a questionnaire-based experiment including 580 participants to explore the impacts of presenting extra information in short videos via text and voice on multi-dimensional user experience. Results indicated that these two elements positively impacted user experience from different dimensions. Gender differences were also found in this study. Based on experimental results, we provided suggestions to support the use of text and voice elements in short video production considering user experience.
Wei-yue Gao, Wei Xiang 0008, Xuanhui Liu, Xueyou Wang, Lingyun Sun
QoMEX2
2019 SmartPaint: a co-creative drawing system based on generative adversarial networks
abstract
Artificial intelligence (AI) has played a significant role in imitating and producing large-scale designs such as e-commerce banners. However, it is less successful at creative and collaborative design outputs. Most humans express their ideas as rough sketches, and lack the professional skills to complete pleasing paintings. Existing AI approaches have failed to convert varied user sketches into artistically beautiful paintings while preserving their semantic concepts. To bridge this gap, we have developed SmartPaint, a co-creative drawing system based on generative adversarial networks (GANs), enabling a machine and a human being to collaborate in cartoon landscape painting. SmartPaint trains a GAN using triples of cartoon images, their corresponding semantic label maps, and edge detection maps. The machine can then simultaneously understand the cartoon style and semantics, along with the spatial relationships among the objects in the landscape images. The trained system receives a sketch as a semantic label map input, and automatically synthesizes its edge map for stable handling of varied sketches. It then outputs a creative and fine painting with the appropriate style corresponding to the human’s sketch. Experiments confirmed that the proposed SmartPaint system successfully generates high-quality cartoon paintings.
Lingyun Sun, Pei Chen 0005, Wei Xiang 0008, Wei-yue Gao
Frontiers Inf. Technol. Electron. Eng.3
2018 Crowdsourcing intelligent design
abstract
Design intelligence, namely, artificial intelligence to solve creative problems and produce creative ideas, has improved rapidly with the new generation artificial intelligence. However, existing methods are more skillful in learning from data and have limitations in creating original ideas different from the training data. Crowdsourcing offers a promising method to produce creative designs by combining human inspiration and machines’ computational ability. We propose a crowdsourcing intelligent design method called ‘flexible crowdsourcing design’. Design ideas produced through crowdsourcing design can be unreliable and inconsistent because they rely solely on selection among participants’ submissions of ideas. In contrast, the flexible crowdsourcing design method employs a cultivation procedure that integrates the ideas from crowd participants and cultivates these ideas to improve design quality at the same time. We introduce a series of studies to show how flexible crowdsourcing design can produce original design ideas consistently. Specifically, we will describe the typical procedure of flexible crowdsourcing design, the refined crowdsourcing tasks, the factors that affect the idea development process, the method for calculating idea development potential, and two applications of the flexible crowdsourcing design method. Finally, it summarizes the design capabilities enabled by crowdsourcing intelligent design. This method enhances the performance of crowdsourcing design and supports the development of design intelligence.
Wei Xiang 0008, Lingyun Sun, Weitao You, Chang-yuan Yang
Frontiers Inf. Technol. Electron. Eng.1
2016 Odor emoticon: An olfactory application that conveys emotions
Wei Xiang 0008, Shi Chen 0005, Lingyun Sun, Shiwei Cheng 0001, V. Michael Bove Jr.
Int. J. Hum. Comput. Stud.1