Yuchong Zhang 0001

dblp:159/8620-1 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-1804-6296ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 12 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AI That Moves With You: A Review of Interactive Technologies Powered by Large Foundation Models for Mobility Impairment
abstract
Large foundation models (FMs) – including large language model (LLM), large vision model (LVM), vision language model (VLM), and related variants – are rapidly reshaping interactive assistive technologies during past years. We present a review of FM-enabled interactive systems for people with mobility impairments, covering work published from January 2020 to May 2025. Searching five databases, we screened 6,249 records and included 26 full papers. We first summerize descriptive results including study design and evaluation approaches of the reviewed studies. We then synthesize FM techniques, model integration patterns, interaction paradigms, and mobility impairment contexts. Our analysis surfaces and distills both technical and ethical challenges existed, lighting up future research topics. We contribute: (i) a conceptualization of FM-enabled interactions for mobility impairment functioning as a design space; (ii) a tabulated corpus with a reproducible codebook; and (iii) a forward agenda to guide and inspire the design of future mobility-assistance interactive systems within human-computer interaction (HCI) and CHI community.
Duosi Dai, Yuchong Zhang 0001, Yong Ma 0003, Danica Kragic
CHI2
2026 Neutral by Default? Replicating User Vocal Responses to Negative Affective Cues in Conversational Agents
abstract
Conversational agents (CAs) increasingly detect users’ emotions, yet deciding how to respond, especially to negative affect, remains a central design challenge. We conducted a role-switching study in which participants reply as the CAs to simulated users expressing anger, sadness, or fear. Results reveal systematic, gender-linked patterns: most male participants favored a neutral, affect-balanced stance and prioritized clarification or task progress, whereas most female participants produced a wider range of non-neutral responses, more often using explicit empathy, reassurance, and reflective listening. We also observe differences in de-escalation phrasing, validation timing, and follow-up questioning across scenarios. These findings indicate that strategies for handling negative emotions vary with user characteristics and context. Based on these findings, we argue for adaptive CA response policies that calibrate first-turn acknowledgment and information-gathering, tailoring prosody and wording to emotional context in order to support de-escalation, perceived understanding, and user trust.
Yong Ma 0003, Yuchong Zhang 0001, Di Fu, Stephanie Zubicueta Portales, Morten Fjeld
HRI2
2026 "Same Voice, Different Language": An Exploration of Voice-Cloned Translation to Support Non-Native Speakers in Online Meetings
abstract
Cross-lingual meetings have become essential for global collaboration, yet current translation technologies often strip away vocal identity — the unique speaker characteristics that convey nuance and social presence. While generic text-to-speech (TTS) provides basic intelligibility, it creates a disconnect between speakers and their translated voices, potentially undermining engagement and comprehension. This paper investigates whether voice cloning technology can bridge this gap by preserving speaker identity in real-time translation. We present a controlled study comparing four voice conditions in meeting interpretation: original speech, gender-neutral TTS, gender-matched TTS, and voice cloning. Through a within-subjects experiment with 45 participants, we demonstrate that voice cloning significantly reduces mental workload (p <.001) and enhances user experience across pragmatic quality (p <.001), hedonic quality (p <.001), and overall satisfaction (p <.001) compared to traditional TTS. While original speech maintained advantages in naturalness, voice cloning achieved superior intelligibility, social impression, and user preference. Qualitative analysis revealed that participants valued voice cloning for preserving speaker identity and improving conversation tracking in multi-speaker scenarios. Our findings suggest that identity-preserving translation represents a significant advancement for cross-lingual communication systems, offering both cognitive and social benefits. We conclude with design implications for integrating voice cloning into meeting platforms while addressing ethical considerations around consent and transparency.
Yong Ma 0003, Yuchong Zhang 0001, Peter Andrews, Zhikun Wu, Stephanie Zubicueta Portales, Morten Fjeld
IUI2
2025 FLAME: A Federated Learning Benchmark for Robotic Manipulation
abstract
Recent progress in robotic manipulation has been fueled by large-scale datasets collected across diverse environments. Training robotic manipulation policies on these datasets is traditionally performed in a centralized manner, raising concerns regarding scalability, adaptability, and data privacy. While federated learning enables decentralized, privacy-preserving training, its application to robotic manipulation remains largely unexplored. We introduce FLAME (Federated Learning Across Manipulation Environments), the first benchmark designed for federated learning in robotic manipulation. FLAME consists of: (i) a set of large-scale datasets of over 160,000 expert demonstrations of multiple manipulation tasks, collected across a wide range of simulated environments; (ii) a training and evaluation framework for robotic policy learning in a federated setting. We evaluate standard federated learning algorithms in FLAME, showing their potential for distributed policy learning and highlighting key challenges. Our benchmark establishes a foundation for scalable, adaptive, and privacy-aware robotic learning. The code is publicly available at https://github.com/KTH-RPL/ELSA-Robotics-Challenge.
Santiago Bou Betran, Alberta Longhini, Miguel Vasco, Yuchong Zhang 0001, Danica Kragic
IROS4
2025 Mind Meets Robots: A Review of EEG-Based Brain-Robot Interaction Systems
abstract
Brain-robot interaction (BRI) empowers individuals to control (semi-)automated machines through brain activity, either passively or actively. In the past decade, BRI systems have advanced significantly, primarily leveraging electroencephalogram (EEG) signals. This article presents an up-to-date review of 87 curated studies published between 2018 and 2023, identifying the research landscape of EEG-based BRI systems. The review consolidates methodologies, interaction modes, application contexts, system evaluation, existing challenges, and future directions in this domain. Based on our analysis, we propose a BRI system model comprising three entities: Brain, Robot, and Interaction, depicting their internal relationships. We especially examine interaction modes between human brains and robots, an aspect not yet fully explored. Within this model, we scrutinize and classify current research, extract insights, highlight challenges, and offer recommendations for future studies. Our findings provide a structured design space for human-robot interaction (HRI), informing the development of more efficient BRI frameworks.
Yuchong Zhang 0001, Nona Rajabi, Farzaneh Taleb, Andrii Matviienko, Yong Ma 0003, Mårten Björkman, Danica Kragic
Int. J. Hum. Comput. Interact.1
2024 Imitation or Innovation? Translating Features of Expressive Motion from Humans to Robots
abstract
Expressive robot motion can help establish acceptance of this technology in everyday life, but understanding what makes movement expressive is a complex and multifaceted task. This paper presents the results of an online study with 46 participants, it aims to explore how people perceive and interpret the expressive qualities of human movement and how they envision the translation of their description into an imagined non-humanoid, quadrupedal robot. Through a qualitative analysis of responses, we conceptualize three themes: their understanding of intent, their interpretations of movement qualities, and finally, their translation from human to robot movement. Respondents’ descriptions of their initial understanding of the performer’s intent fall into two modes, bio-mechanical and narrative. We illustrate their interpretations of movement qualities through four strategies: movement features as kinematic indicators, intent indicators, attributed context, and perceived internal states. Lastly, we observe their translation from human to robot movement, with a particular focus on respondents’ use of kinaesthetic empathy and anthropomorphism. Our findings aim to support a bottom-up approach, using users’ general knowledge for designing expressive robot motion.
Benedikte Wallace, Marieke van Otterdijk, Yuchong Zhang 0001, Nona Rajabi, Diego Marin-Bucio, Danica Kragic, Jim Tørresen
HAI3
2024 Human-centered AI Technologies in Human-robot Interaction for Social Settings
abstract
The increasing integration of human-robot interaction (HRI) into social settings demands the development of human-centered AI technologies that prioritize intuitive, ethical, and empathetic interactions. As robots become more prevalent in everyday life—ranging from assistive devices in healthcare to educational tools in classrooms and customer service agents in retail—it is essential to ensure they can communicate and collaborate with humans in ways that are not only effective but also socially appropriate and meaningful. This workshop aims to explore cutting-edge advancements and interdisciplinary approaches to building AI-driven systems that facilitate effective, meaningful, and socially appropriate interactions between robots and humans across various environments such as healthcare, education, and customer service. We will primarily focus on several key themes, such as human-centered contextual AI, AI-driven intelligent robotics, ethical and responsible AI, and real-world applications. This workshop invites contributions from researchers, practitioners, and developers who are working on AI systems that empower robots to operate effectively in human-centered environments. By addressing challenges such as interpreting human emotions, understanding social cues, and adhering to ethical standards, and by sharing advancements in human-centered AI, we aim to shape the future of HRI. Our goal is to ensure that robots enrich human social experiences, fostering interactions that are not only efficient but also enhance the quality of life. By uniting efforts from various disciplines, we aspire to create robots that seamlessly integrate into society, ultimately contributing to a more harmonious coexistence between humans and robotic systems.
Yuchong Zhang 0001, Khaled Kassem, Zhengya Gong, Yong Ma 0003, Emma Kirjavainen, Jonna Häkkilä
MUM1
2023 Playing with Data: An Augmented Reality Approach to Interact with Visualizations of Industrial Process Tomography
Yuchong Zhang 0001, Yueming Xuan, Adel Omrani, Morten Fjeld
INTERACT (2)1
2023 See or Hear? Exploring the Effect of Visual/Audio Hints and Gaze-assisted Instant Post-task Feedback for Visual Search Tasks in AR
abstract
Augmented reality (AR) is emerging in visual search tasks for increasingly immersive interactions with virtual objects. We propose an AR approach providing visual and audio hints along with gaze-assisted instant post-task feedback for search tasks based on mobile head-mounted display (HMD). The target case was a book-searching task, in which we aimed to explore the effect of the hints together with the task feedback with two hypotheses. H1: Since visual and audio hints can positively affect AR search tasks, the combination outperforms the individuals. H2: The gaze-assisted instant post-task feedback can positively affect AR search tasks. The proof-of-concept was demonstrated by an AR app in HMD and a comprehensive user study (n=96) consisting of two sub-studies, Study I (n=48) without task feedback and Study II (n=48) with task feedback. Following quantitative and qualitative analysis, our results partially verified H1 and completely verified H2, enabling us to conclude that the synthesis of visual and audio hints conditionally improves the AR visual search task efficiency when coupled with task feedback.
Yuchong Zhang 0001, Adam Nowak, Yueming Xuan, Andrzej Romanowski, Morten Fjeld
ISMAR1
2022 On-site or Remote Working?: An Initial Solution on How COVID-19 Pandemic May Impact Augmented Reality Users
abstract
As a cutting edge technique requiring high-precision equipment, augmented reality (AR) and its users are influenced by the ambient environment. With the tremendous effect brought by COVID-19 pandemic, most people have shifted from on-site working to remote working. In this study, we propose an initial solution to explore the impact of COVID-19 pandemic on AR users working in these two situations. We develop a prototype application facilitated with gamification process in which users are requested to play an AR game in headset both in on-site and remote working environments. This game, which is highly dependent on the ambient environment, enables people to memorize, distinguish, and place virtual objects when immersing themselves into different surroundings with distinct distractors. We envision to conduct more user studies investigating how COVID-19 affects AR users, which could lead to more in-depth studies in the future.
Yuchong Zhang 0001, Adam Nowak, Andrzej Romanowski, Morten Fjeld
AVI1
2022 An Initial Exploration of Visual Cues in Head-Mounted Display Augmented Reality for Book Searching
abstract
Augmented reality (AR) is today becoming more widely utilized as it allows for interacting with the virtual objects. In this study, we propose an head-mounted display (HMD) AR system supporting book searching with visual cues. The visual cue is represented as a light-green blob hinting users for task completion, which significantly strengthens the overall performance. The system is implemented by using Microsoft HoloLens 2. The proof-of-concept version of the proposed solution is demonstrated in a pilot user study (n=8) comprising an experimental group (with visual cues, n=4) and a control group (without visual cues, n=4), followed by quantitative analysis of task completion time (TCT) and NASA task load index (TLX). The results show that our proposed HMD AR solution improved the task performance and had cognitive benefits for book searching tasks.
Yuchong Zhang 0001, Adam Nowak, Andrzej Romanowski, Morten Fjeld
MUM1
2019 The Design of Social Drones: A Review of Studies on Autonomous Flyers in Inhabited Environments
abstract
The design space of social drones, where autonomous flyers operate in close proximity to human users or bystanders, is distinct from use cases involving a remote human operator and/or an uninhabited environment; and warrants foregrounding human-centered design concerns. Recently, research on social drones has followed a trend of rapid growth. This paper consolidates the current state of the art in human-centered design knowledge about social drones through a review of relevant studies, scaffolded by a descriptive framework of design knowledge creation. Our analysis identified three high-level themes that sketch out knowledge clusters in the literature, and twelve design concerns which unpack how various dimensions of drone aesthetics and behavior relate to pertinent human responses. These results have the potential to inform and expedite future research and practice, by supporting readers in defining and situating their future contributions. The materials and results of our analysis are also published in an open online repository that intends to serve as a living hub for a community of researchers and designers working with social drones.
Mehmet Aydin Baytas, Damla Çay, Yuchong Zhang 0001, Mohammad Obaid, Asim Evren Yantaç, Morten Fjeld
CHI3
2019 Automatic Image Segmentation for Microwave Tomography (MWT): From Implementation to Comparative Evaluation
abstract
Inspired by its high performance in image-based medical analysis, this poster paper explores the use of advanced segmentation techniques for industrial Microwave Tomography (MWT). Our context is the visual analysis of moisture levels in porous foams undergoing microwave drying. We propose an automatic segmentation technique---MWT Segmentation based on K-means (MWTS-KM) and demonstrate its efficiency and accuracy for industrial use. MWTS-KM consists of three stages: image augmentation, grey-scale conversion, and K-means implementation. To estimate the performance of this technique, we empirically benchmark its efficiency and accuracy against two well-established alternatives: Otsu and K-means. To elicit performance data, three metrics (Jaccard index, Dice coefficient and false positive) are used. Based on our experiments, our results indicate that MWTS-KM outperforms the well-established Otsu and K-means.
Yuchong Zhang 0001, Yong Ma 0003, Adel Omrani, Morten Fjeld, Marco Fratarcangeli
VINCI1