VLDB 2026 Research / reviewers in the wild / expert
Shi Chen 0005
dblp:24/2311-5
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-3577-5725ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round DebateabstractAutomating Graphical User Interface (GUI) operations with Multimodal Large Language Models (MLLMs) is promising but remains bottlenecked in real-world long-horizon settings. Key challenges include ensuring precise grounding across diverse interfaces and handling irreversible errors in extended workflows. Current methods often struggle to distinguish targets in low Signal-to-Noise Ratio (SNR) environments and lack sufficient pre-execution verification to prevent error accumulation. To address this, we propose the Memory-augmented Debate System (MaDS). Specifically, MaDS combines: (1) a Dual-Layer Memory Module that integrates universal interaction priors with scenario-specific operational experience to mitigate grounding hallucinations; and (2) Multi-Round Debate that performs pre-execution verification, while transforming execution failures into retrievable Negative Warnings to reduce repeated errors. Additionally, we introduce MaDS-Benchmark, a benchmark for long-horizon mobile GUI tasks with process-oriented evaluation. Experiments show that MaDS achieves a 90.23% Task Success Rate on MaDS-Benchmark and strong performance on public benchmarks including AITW, AITZ, CAGUI, and GUIOdyssey. Pengchen Chen, Shi Chen 0005, Qiming Ye, Xinli Chen, Wei Xiang 0008 |
ACL (1) | 2 |
| 2026 | Pika: Designing a social-support agent to improve drivers' experience in gig work
Wei Xiang 0008, Xinli Chen, Tianhui Guo, Shi Chen 0005 |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | Voice by the Non-sighted: Practices and Challenges of Audiobook Voice Actors with Blind and Low Vision in China
Shi Chen 0005, Jingao Zhang, Suqi Lou, Wei Xiang 0008, Lingyun Sun |
CHI | 1 |
| 2025 | VAEnvGen: A Real-Time Virtual Agent Environment Generation System Based on Large Language ModelsabstractEnvironment plays an important role in non-verbal communication for human-virtual agent interaction. Existing research explores the influence of an agent’s appearance and attributes to enhance human-virtual agent communication. However, there is no common practice for dynamically adjusting the surrounding environments of the virtual agent. In this paper, we introduce a real-time virtual agent environment generation system (VAEnvGen), which contributes to the field by enhancing users’ content perception and improving task performance through dynamic environment adjustment. The system dynamically analyzes both the appropriate communication environment and filters the key information according to the current context. Leveraging Large Language Models, it generates a pseudo-3D background space to create an engaging atmosphere and a dynamic foreground content space for vivid key information display, thereby significantly enhancing content perception. For widespread adoption and flexibility, VAEnvGen is developed as a web application. We further evaluate the impact of VAEnvGen on content perception, user attention, and subjective satisfaction through a mixed-design user study with 50 participants. Quantitative and qualitative results reveal significant improvements in content perception, task completion time, and user satisfaction when using VAEnvGen. The system effectively redistributes user attention from subtitles and the virtual agent itself to the dynamically generated background and key foreground information, leading to a more immersive and less fatiguing user experience. Jingyu Wu, Pengchen Chen, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun |
Int. J. Hum. Comput. Interact. | 3 |
| 2025 | Behind the Same Mask: Understanding the Practice of Spontaneous Collective Anonymity on Chinese Social PlatformsabstractAnonymity plays a crucial role in social interactions online. Recently, a new phenomenon has emerged on Chinese social platforms where users collectively adopt a uniform avatar and nickname ''momo'', thereby achieving anonymity. However, understanding such spontaneous collective anonymity within Chinese cultural and contextual factors remains limited since much of the anonymity research focuses on Western users. Yet, it is unclear how users perceive the usage of ''momo'', their motivations, and how using this collective anonymity impacts their social interaction. To answer these questions, we conducted interviews with 20 ''momo'' users. We found that the shared identity ''momo'' provides an additional layer of anonymity on identity-constrained Chinese social platforms. Users adopted ''momo'' to engage in more inclusive discussions and to balance anonymity and self-presentation. Moreover, this collective anonymity fosters connections and forms a meaningful group identity in a loosely organized community. We also identified the benefits and risks associated with this unique collective anonymity. This work makes significant contributions to CSCW and HCI research by (1) extending the knowledge of anonymity practices and privacy concerns within non-Western and mainly Chinese contexts. (2) advancing the work on anonymity models by revealing the dual role of the Momo identity in facilitating collective anonymity and community bonds. (3) providing design implications to support future social technologies in identity design and anonymous communities. Suqi Lou, Chao Zhang 0082, Shi Chen 0005, Zhicong Lu, Yaxing Yao |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2024 | Silent Delivery: Practices and Challenges of Delivering Among Deaf or Hard of Hearing CouriersabstractThis paper explores the motivations, practices, and challenges of Deaf and Hard of Hearing (DHH) couriers in China’s food delivery industry. Interviews reveal a preference for this industry due to better pay, job satisfaction, and community belonging. DHH couriers tend to and frequently disclose their DHH disability using platform tags and text messages. They also utilize accessible communication tools provided by the delivery platforms, such as AI voice calls, voice-to-text technologies, and electronic communication cards, to facilitate communication during the delivery process. Despite these technological aids, human intervention remains crucial throughout the delivery process. Challenges encountered include safety risks when riding mopeds, the complexities of multitasking, and user mistrust in AI voice calls. Our findings offer valuable insights for designing more inclusive delivery platforms and have broader implications for creating employment opportunities for DHH, particularly in developing countries. Shi Chen 0005, Jingao Zhang, Yuge Qi, Jiaqi Teng, Zhihan Zeng |
CHI | 1 |
| 2024 | SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile ApplicationsabstractThe conflict between the rapid iteration demand of prototyping and the time-consuming nature of user tests has led researchers to adopt AI methods to identify usability issues. However, these AI-driven methods concentrate on evaluating the feasibility of a system, while often overlooking the influence of specified user characteristics and usage contexts. Our work proposes a tool named SimUser based on large language models (LLMs) with the Chain-of-Thought structure and user modeling method. It generates usability feedback by simulating the interaction between users and applications, which is influenced by user characteristics and contextual factors. The empirical study (48 human users and 21 designers) validated that in the context of a simple smartwatch interface, SimUser could generate heuristic usability feedback with the similarity varying from 35.7% to 100% according to the user groups and usability category. Our work provides insights into simulating users by LLM to improve future design activities. Wei Xiang 0008, Hanfei Zhu, Suqi Lou, Xinli Chen, Zhenghua Pan, Yuping Jin, Shi Chen 0005, Lingyun Sun |
CHI | 7 |
| 2024 | Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed InputsabstractMulti-modal Large Language Models (MLLMs) have demonstrated remarkable performance on various visual-language understanding and generation tasks. However, MLLMs occasionally generate content inconsistent with the given images, which is known as "hallucination". Prior works primarily center on evaluating hallucination using standard, unperturbed benchmarks, which overlook the prevalent occurrence of perturbed inputs in real-world scenarios-such as image cropping or blurring-that are critical for a comprehensive assessment of MLLMs' hallucination. In this paper, to bridge this gap, we propose Hallu-PI, the first benchmark designed to evaluate Hallucination in MLLMs within Perturbed Inputs. Specifically, Hallu-PI consists of seven perturbed scenarios, containing 1,260 perturbed images from 11 object types. Each image is accompanied by detailed annotations, which include fine-grained hallucination types, such as existence, attribute, and relation. We equip these annotations with a rich set of questions, making Hallu-PI suitable for both discriminative and generative tasks. Extensive experiments on 12 mainstream MLLMs, such as GPT-4V and Gemini-Pro Vision, demonstrate that these models exhibit significant hallucinations on Hallu-PI, which is not observed in unperturbed scenarios. Furthermore, our research reveals a severe bias in MLLMs' ability to handle different types of hallucinations. We also design two baselines specifically for perturbed scenarios, namely Perturbed-Reminder and Perturbed-ICL. We hope that our study will bring researchers' attention to the limitations of MLLMs when dealing with perturbed inputs, and spur further investigations to address this issue. Our code and datasets are publicly available at https://github.com/NJUNLP/Hallu-PI. Peng Ding 0001, Jingyu Wu, Jun Kuang, Dan Ma 0008, Xuezhi Cao, Shi Chen 0005, Jiajun Chen 0001, Shujian Huang |
ACM Multimedia | 7 |
| 2024 | CNAMD Corpus: A Chinese Natural Audiovisual Multimodal Database of Conversations for Social Interactive AgentsabstractImpressive progress has been made in developing companion Socially Interactive Agents (SIAs) that provide companionship and reduce loneliness. However, recent works focus on analyzing multimodal feedback in Answer part but ignore Question part. Furthermore, research on SIAs is primarily based on English, which poses a challenge for Chinese SIAs because of cultural differences between English and Chinese. Therefore, we introduce a Chinese Natural Audiovisual Multimodal Database (CNAMD) corpus, the first and largest freely available Chinese multimodal database for multi-person interaction, containing 48 hours of videos and annotations across eight modalities. Using CNAMD, we analyze the characteristics of vocal-verbal, audio, behavioral, and multimodal combinations during questioning, test the performance of six baselines on three tasks, and propose improvements for processing daily Chinese data. The present findings will help designers consider Chinese customs and language when designing Chinese SIAs, making them more suitable for the Chinese cultural context and users. Jingyu Wu, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun, Hongzeng Zhang, Yanxu Li |
Int. J. Hum. Comput. Interact. | 2 |
| 2024 | Automatic Generation of Interactive Nonlinear Video for Online Apparel Shopping NavigationabstractWe present an automatic generation pipeline of interactive nonlinear video for online apparel shopping navigation. Our approach was inspired by Google's “Messy Middle” theory, which suggests that people mentally are faced with two tasks—exploration and evaluation—before purchasing online. Given a set of apparel product presentation videos, our navigation UI organizes them to optimize users' product exploration and automatically generates interactive videos for users' product evaluation. To support automatic methods, we proposed a video clustering similarity ($\operatorname{CSIM}$) and a camera movement similarity ($\operatorname{MSIM}$), as well as a comparative video generation algorithm for product recommendation, presentation, and comparison. To evaluate our pipeline's effectiveness, we conducted several user studies. The results showed that our pipeline can help users complete the consumption process more efficiently, making it easier for them to understand and choose a product. Weitao You, Juntao Ji, Lingyun Sun, Chang-yuan Yang, Mi Yu, Shi Chen 0005 |
IEEE Trans. Multim. | 6 |
| 2023 | "I Never Envy Anyone, for I Have Already Built a Kingdom With My Fingertips": Exploring Teenagers' Experience in Chat-based Cosplay CommunityabstractThis paper reports an interview study about the practice of teenagers’ chat-based cosplay in China. Findings reveal the four primary motivations of the participants and their main practice in chat-based cosplay. We found that adolescents perceived character presentation and portrayal as a central aspect of chat-based cosplay and they devoted significant effort to refine their characters to achieve higher character consistency. We highlighted the positive feedback loop between social relationships and story creation in chat-based cosplay community. In addition, we identified the influence and negative experiences on adolescents in the chat-based cosplay community. Yaohua Bu, Suqi Lou, Shi Chen 0005, Lingyun Sun, Chang-yuan Yang |
IDC | 4 |
| 2023 | Cultural Self-Adaptive Multimodal Gesture Generation Based on Multiple Culture Gesture DatasetabstractCo-speech gesture generation is essential for multimodal chatbots and agents. Previous research extensively studies the relationship between text, audio, and gesture. Meanwhile, to enhance cross-culture communication, culture-specific gestures are crucial for chatbots to learn cultural differences and incorporate cultural cues. However, culture-specific gesture generation faces two challenges: lack of large-scale, high-quality gesture datasets that include diverse cultural groups, and lack of generalization across different cultures. Therefore, in this paper, we first introduce a Multiple Culture Gesture Dataset (MCGD), the largest freely available gesture dataset to date. It consists of ten different cultures, over 200 speakers, and 10,000 segmented sequences. We further propose a Cultural Self-adaptive Gesture Generation Network (CSGN) that takes multimodal relationships into consideration while generating gestures using a cascade architecture and learnable dynamic weight. The CSGN adaptively generates gestures with different cultural characteristics without the need to retrain a new network. It extracts cultural features from the multimodal inputs or a cultural style embedding space with a designated culture. We broadly evaluate our method across four large-scale benchmark datasets. Empirical results show that our method achieves multiple cultural gesture generation and improves comprehensiveness of multimodal inputs. Our method improves the state-of-the-art average FGD from 53.7 to 48.0 and culture deception rate (CDR) from 33.63% to 39.87%. Jingyu Wu, Shi Chen 0005, Shuyu Gan, Chang-yuan Yang, Lingyun Sun |
ACM Multimedia | 2 |
| 2023 | What makes virtual intimacy...intimate? Understanding the Phenomenon and Practice of Computer-Mediated Paid CompanionshipabstractVirtual romance service (VRS), as a notable commodification of intimacy, is currently emerging in China. Such service is not similar to the kind of intimacy that fans and idols generate through parasocial relationships, but behaves as the direct dyadic intimacy between service providers (virtual lovers) and buyers (customers). To gain a deep understanding of computer-mediated paid companionship, we study emerging user behaviors in VRS through a mixed-method study, including a survey (N = 178) and a follow-up semi-structured interview (N = 22) with both virtual lovers and customers to learn about their motivations, perceptions, and how virtual lovers provide online paid companionship to meet customers' emotional needs. We found three behavioral strategies of virtual lovers and the fact that they provide service in surface and deep acting and real feeling. Customers see VRS as a way to obtain affective benefits with reduced affective cost. We also found that VRS customers paid for the tangible benefits of an idealized romantic partner, rather than long-term commitment and emotional investment, and we identified key characteristics that VRS reduces from intimate relationships that fit its pay-per-use feature. We conclude by discussing the nature of virtual lovers and design implications for computer-mediated paid companionship. Shi Chen 0005, Lingyun Sun, Chang-yuan Yang |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2016 | Odor emoticon: An olfactory application that conveys emotions
Wei Xiang 0008, Shi Chen 0005, Lingyun Sun, Shiwei Cheng 0001, V. Michael Bove Jr. |
Int. J. Hum. Comput. Stud. | 2 |
| 2015 | Emotion-driven Chinese folk music-image retrieval based on DE-SVM
Baixi Xing, Shouqian Sun, Lekai Zhang, Zenggui Gao, Shi Chen 0005 |
Neurocomputing | 7 |