Xin Sun 0016

dblp:20/3535-16 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-8188-7576ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
abstract
Large language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased by disclosed source labels. Using a counterfactual design, we find that both humans and LLM judges assign higher trust to information labeled as human-authored than to the same content labeled as AI-generated. Eye-tracking data reveal that humans rely heavily on source labels as heuristic cues for judgments. We analyze LLM internal states during judgment. Across label conditions, models allocate denser attention to the label region than the content region, and this label dominance is stronger under Human labels than AI labels, consistent with the human gaze patterns. Besides, decision uncertainty measured by logits is higher under AI labels than Human labels. These results indicate that the source label is a salient heuristic cue for both humans and LLMs. It raises validity concerns for label-sensitive LLM-as-a-Judge evaluation, and we cautiously raise that aligning models with human preferences may propagate human heuristic reliance into models, motivating debiased evaluation and alignment.
Xin Sun 0016, Sijing Qin, Isao Echizen, Abdallah El Ali, Saku Sugawara
ACL (1)1
2026 Zenergy: Designing Taoist-Inspired Transformative Nature Imagery for Everyday Empowerment
abstract
Nature has long been valued for its restorative impact on emotion and well-being, motivating many HCI systems to incorporate nature as a calming design material. However, cultural traditions such as Taoism frame nature not as passive, but as an active, symbolic force for emotional transformation. We present Zenergy, a mobile application that uses a large language model to generate personalized guided meditations grounded in Taoist nature imagery. Based on users’ emotional and contextual input, Zenergy leads them through a symbolic journey using natural metaphors such as wind to release burdens, rivers to restore flow, and sunlight to renew strength. A mixed-method field study (N = 27) showed that Zenergy enhanced users’ self-efficacy, emotional clarity, and spiritual connection. We introduce transformative nature imagery as a design lens for everyday empowerment, and offer strategies for embedding culturally grounded symbolism into interactive well-being technologies.
Zhuying Li 0001, Yishu Wang 0010, Yan Wang 0057, Xin Sun 0016
CHI4
2026 Eyes Can't Always Tell: Fusing Eye Tracking and User Priors for User Modeling under AI Advice Conditions
abstract
Modeling users' cognitive states (e.g., cognitive load and decision confidence) is essential for building adaptive AI in high-stakes decision-making. While eye tracking provides non-invasive behavioral signals correlated with cognitive effort, prior work has not systematically examined how AI assistance contexts, specifically varying advice reliability and user heterogeneity, can alter the mapping between gaze signals and cognitive states. We conducted a within-subject lab eye-tracking study (N=54) on factual verification tasks under three conditions: No-AI, Correct-AI advice, and Incorrect-AI advice. We analyze condition-dependent changes in self-reports and eye-tracking patterns and evaluate the robustness of eye-tracking-based user modeling. Results show that AI advice increases decision confidence compared to No-AI, while Correct-AI is associated with lower perceived cognitive load and more efficient gaze behavior. Crucially, predictive modeling is context-sensitive: the relationship between eye-tracking signals and cognitive states shifts across AI conditions. Finally, fusing eye-tracking features with user priors (demographics, AI literacy/experience, and propensity to trust technology) improves cross-participant generalization. These findings support condition-aware and personalized user modeling for cognitively aligned adaptive AI systems.
Xin Sun 0016, Shu Wei, Jos A. Bosch, Isao Echizen, Abdallah El Ali, Saku Sugawara
UMAP1
2026 Understanding trust toward human versus AI-generated health information through behavioral and physiological sensing
abstract
As AI-generated health information proliferates online and becomes increasingly indistinguishable from human-sourced information, it becomes critical to understand how people trust and label such content, especially when the information is inaccurate. We conducted two complementary studies: (1) a mixed-methods survey (N=142) employing a 2 (source: Human vs. LLM) × 2 (label: Human vs. AI) × 3 (type: General, Symptom, Treatment) design, and (2) a within-subjects lab study (N=40) incorporating eye-tracking and physiological sensing (ECG, EDA, skin temperature). Participants were presented with health information varying by source-label combinations and asked to rate their trust, while their gaze behavior and physiological signals were recorded. We found that LLM-generated information was trusted more than human-generated content, whereas information labeled as human was trusted more than that labeled as AI. Trust remained consistent across information types. Eye-tracking and physiological responses varied significantly by source and label. Machine learning models trained on these behavioral and physiological features predicted binary self-reported trust levels with 73 % accuracy and information source with 65 % accuracy. Our findings demonstrate that adding transparency labels to online health information modulates trust. Behavioral and physiological features show potential to verify trust perceptions and indicate if additional transparency is needed.
Xin Sun 0016, Rongjun Ma, Shu Wei, Pablo César, Jos A. Bosch, Abdallah El Ali
Int. J. Hum. Comput. Stud.1
2026 Digital Spirituality in Mainland China: Understanding Online Practices for Designing Culturally Relevant Spiritual Experiences CSCW013
abstract
Spiritual engagement is increasingly mediated by digital platforms, yet little is known about how this shift unfolds in non-Western contexts. In mainland China, individuals are turning to online social platforms like Douyin, WeChat, and Xiaohongshu to access spiritual content that blends religious, cultural, and philosophical traditions. In this paper, we present a mixed-methods study that combines a survey (N = 207) with in-depth interviews (N = 16) to examine how users engage with online spiritual content. The survey outlines general usage patterns across platforms, content types, and motivations, while the interviews reveal four key themes: the pragmatic use of spirituality to navigate emotional and material challenges; the flexible spiritual engagement; a cultural reconnection with spiritual heritage; and concerns over trust, authenticity, and regulation in the digital spiritual landscape. Based on these findings, we offer design implications for creating culturally grounded online spiritual experiences and contribute to HCI literature by advancing understanding of how sociotechnical systems can support spiritually meaningful engagement in diverse cultural contexts.
Zhuying Li 0001, Ziteng Zhang, Xin Sun 0016, Yan Wang 0057
Proc. ACM Hum. Comput. Interact.3
2025 Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
abstract
Kawaii" is the Japanese concept of cute, which carries sociocultural connotations related to social identities and emotional responses.Yet, virtually all work to date has focused on the visual side of kawaii, including in studies of computer agents and social robots.In pursuit of formalizing the new science of kawaii vocalics, we explored what elements of voice relate to kawaii and how they might be manipulated, manually and automatically.We conducted a four-phase study (grand 𝑁 = 512) with two varieties of computer voices: text-to-speech (TTS) and game character voices.We found kawaii "sweet spots" through manipulation of fundamental and formant frequencies, but only for certain voices and to a certain extent.Findings also suggest a ceiling effect for the kawaii vocalics of certain voices.We offer empirical validation of the preliminary kawaii vocalics model and an elementary method for manipulating kawaii perceptions of computer voice.
Yuto Mandai, Katie Seaborn, Tomoyasu Nakano, Xin Sun 0016, Yijia Wang 0001, Jun Kato 0001
CHI4
2025 How Well Can Large Language Models Reflect? A Human Evaluation of LLM-generated Reflections for Motivational Interviewing Dialogues
abstract
Motivational Interviewing (MI) is a counseling technique that promotes behavioral change through reflective responses to mirror or refine client statements. While advanced Large Language Models (LLMs) can generate engaging dialogues, challenges remain for applying them in a sensitive context such as MI. This work assesses the potential of LLMs to generate MI reflections via three LLMs: GPT-4, Llama-2, and BLOOM, and explores the effect of dialogue context size and integration of MI strategies for reflection generation by LLMs. We conduct evaluations using both automatic metrics and human judges on four criteria: appropriateness, relevance, engagement, and naturalness, to assess whether these LLMs can accurately generate the nuanced therapeutic communication required in MI. While we demonstrate LLMs’ potential in generating MI reflections comparable to human therapists, content analysis shows that significant challenges remain. By identifying the strengths and limitations of LLMs in generating empathetic and contextually appropriate reflections in MI, this work contributes to the ongoing dialogue in enhancing LLM’s role in therapeutic counseling.
Mustafa Erkan Basar, Xin Sun 0016, Iris Hendrickx, Jan de Wit, Tibor Bosse, Gert-Jan de Bruijn, Jos A. Bosch, Emiel Krahmer
COLING2
2025 Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies
abstract
Recent advancements in large language models (LLMs) have shown promise in generating psychotherapeutic dialogues, particularly in the context of motivational interviewing (MI). However, the inherent lack of transparency in LLM outputs presents significant challenges given the sensitive nature of psychotherapy. Applying MI strategies, a set of MI skills, to generate more controllable therapeutic-adherent conversations with explainability provides a possible solution. In this work, we explore the alignment of LLMs with MI strategies by first prompting the LLMs to predict the appropriate strategies as reasoning and then utilizing these strategies to guide the subsequent dialogue generation. We seek to investigate whether such alignment leads to more controllable and explainable generations. Multiple experiments including automatic and human evaluations are conducted to validate the effectiveness of MI strategies in aligning psychotherapy dialogue generation. Our findings demonstrate the potential of LLMs in producing strategically aligned dialogues and suggest directions for practical applications in psychotherapeutic settings.
Xin Sun 0016, Abdallah El Ali, Zhuying Li 0001, Pengjie Ren, Jan de Wit, Jiahuan Pei, Jos A. Bosch
COLING1
2025 Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
abstract
With large language models (LLMs) on the rise, in-game interactions are shifting from rigid commands to natural conversations. However, the impacts of LLMs on player performance and game experience remain underexplored. This work explores LLM's role as a co-builder during gameplay, examining its impact on task performance, usability, and player experience. Using Minecraft as a sandbox, we present an LLM-assisted interface that engages players through natural language, aiming to facilitate creativity and simplify complex gaming commands. We conducted a mixed-methods study with 30 participants, comparing LLM-assisted and command-based interfaces across simple and complex game tasks. Quantitative and qualitative analyses reveal that the LLM-assisted interface significantly improves player performance, engagement, and overall game experience. Additionally, task complexity has a notable effect on player performance and experience across both interfaces. Our findings highlight the potential of LLM-assisted interfaces to revolutionize virtual experiences, emphasizing the importance of balancing intuitiveness with predictability, transparency, and user agency in AI-driven, multimodal gaming environments.
Xin Sun 0016, Yue Li 0044, Jie Li 0064, Massimo Poesio, Julian Frommel, Koen V. Hindriks, Jiahuan Pei
ICMI1
2025 MultiJustice: A Chinese Dataset for Multi-party, Multi-charge Legal Prediction
Jiahuan Pei, Diancheng Shui, Zhiguang Han, Xin Sun 0016
NLPCC (1)5
2025 Integrating culture in Human-Food Interaction: A study of cultural and creative food experiences and technological interactions
Zhuying Li 0001, Yan Wang 0057, Xin Sun 0016
Int. J. Hum. Comput. Stud.3
2025 Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for Psychotherapy
abstract
Chatbots or conversational agents (CAs) are increasingly used to improve access to digital psychotherapy. Many current systems rely on rigid, rule-based designs, heavily dependent on expert-crafted dialogue scripts for guiding therapeutic conversations. Although advances in large language models (LLMs) offer potential for more flexible interactions, their lack of controllability and explanability poses challenges in high-stakes contexts like psychotherapy. To address this, we conducted two studies in this work to explore how aligning LLMs with expert-crafted scripts can enhance psychotherapeutic chatbot performance. In Study 1 (N=43), an online experiment with a within-subjects design, we compared rule-based, pure LLM, and LLMs aligned with expert-crafted scripts via fine-tuning and prompting. Results showed that aligned LLMs significantly outperformed the other types of chatbots in empathy, dialogue relevance, and adherence to therapeutic principles. Building on findings, we proposed ''Script-Strategy Aligned Generation (SSAG)'', a more flexible alignment approach that reduces reliance on fully scripted content while maintaining LLMs' therapeutic adherence and controllability. In a 10-day field Study 2 (N=21), SSAG achieved comparable therapeutic effectiveness to full-scripted LLMs while requiring less than 40% of expert-crafted dialogue content. Beyond these results, this work advances LLM applications in psychotherapy by providing a controllable and scalable solution, reducing reliance on expert effort. By enabling domain experts to align LLMs through high-level strategies rather than full scripts, SSAG supports more efficient co-development and expands access to a broader context of psychotherapy.
Xin Sun 0016, Jan de Wit, Zhuying Li 0001, Jiahuan Pei, Abdallah El Ali, Jos A. Bosch
Proc. ACM Hum. Comput. Interact.1
2025 Interface Matters: Exploring Human Trust in Health Information from Large Language Models via Text, Speech, and Embodiment
abstract
The deployment of Conversational User Interfaces (CUIs) with advanced Large Language Models (LLMs) has significantly transformed health information seeking and dissemination, facilitating immediate and interactive communication between users and digital health resources. However, while trust is crucial for adopting health advice, how the dissemination interface influences people's perceived trust in health information provided by LLMs remains unclear. To address this, we conducted a mixed-methods, within-subjects lab study (N=20) to investigate how different CUIs (i.e., a text-based, speech-based, and embodied interface) affect user-perceived trust levels when delivering health information from an identical LLM source. Our key findings showed that: (a) participants' trust levels in health information delivered were significantly variant across different interfaces; (b) there are significant correlations between trust in health-related information and trust in the delivered interface as well as the usability level of the interface; (c) the type of health questions did not affect participants' perceived trust. Besides, we identified key factors influencing trust in health information delivered through various CUIs and explored differences in how people trust health information from LLM and its dissemination. We highlight the potential of LLM-powered CUIs in supporting health-related information-seeking behaviors. This work contributes insights for ensuring effective and trustworthy personal health information-seeking in the era of LLM-powered CUIs and multi-modal information dissemination.
Xin Sun 0016, Jos A. Bosch, Zhuying Li 0001
Proc. ACM Hum. Comput. Interact.1
2024 Eliciting Motivational Interviewing Skill Codes in Psychotherapy with LLMs: A Bilingual Dataset and Analytical Study
abstract
Behavioral coding (BC) in motivational interviewing (MI) holds great potential for enhancing the efficacy of MI counseling. However, manual coding is labor-intensive, and automation efforts are hindered by the lack of data due to the privacy of psychotherapy. To address these challenges, we introduce BiMISC, a bilingual dataset of MI conversations in English and Dutch, sourced from real counseling sessions. Expert annotations in BiMISC adhere strictly to the motivational interviewing skills code (MISC) scheme, offering a pivotal resource for MI research. Additionally, we present a novel approach to elicit the MISC expertise from Large language models (LLMs) for MI coding. Through the in-depth analysis of BiMISC and the evaluation of our proposed approach, we demonstrate that the LLM-based approach yields results closely aligned with expert annotations and maintains consistent performance across different languages. Our contributions not only furnish the MI community with a valuable bilingual dataset but also spotlight the potential of LLMs in MI coding, laying the foundation for future MI research.
Xin Sun 0016, Jiahuan Pei, Jan de Wit, Mohammad Aliannejadi, Emiel Krahmer, Jos T. P. Dobber, Jos A. Bosch
LREC/COLING1