VLDB 2026 Research / reviewers in the wild / expert
Yong Ma 0003
dblp:33/3013-3
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-8398-4118ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI That Moves With You: A Review of Interactive Technologies Powered by Large Foundation Models for Mobility ImpairmentabstractLarge foundation models (FMs) – including large language model (LLM), large vision model (LVM), vision language model (VLM), and related variants – are rapidly reshaping interactive assistive technologies during past years. We present a review of FM-enabled interactive systems for people with mobility impairments, covering work published from January 2020 to May 2025. Searching five databases, we screened 6,249 records and included 26 full papers. We first summerize descriptive results including study design and evaluation approaches of the reviewed studies. We then synthesize FM techniques, model integration patterns, interaction paradigms, and mobility impairment contexts. Our analysis surfaces and distills both technical and ethical challenges existed, lighting up future research topics. We contribute: (i) a conceptualization of FM-enabled interactions for mobility impairment functioning as a design space; (ii) a tabulated corpus with a reproducible codebook; and (iii) a forward agenda to guide and inspire the design of future mobility-assistance interactive systems within human-computer interaction (HCI) and CHI community. Duosi Dai, Yuchong Zhang 0001, Yong Ma 0003, Danica Kragic |
CHI | 3 |
| 2026 | Beyond Words: Measuring User Experience through Speech Analysis in Voice User InterfacesabstractVoice assistants (VAs) are typically evaluated through task performance metrics and self-report questionnaires, but people’s voices themselves carry rich paralinguistic cues that reveal affect, effort, and interaction breakdowns. We present a within-subjects study (N=49) that systematically compared three VA personas across three usage scenarios to investigate whether speech-derived audio features can serve as a proxy for user experience (UX). Participants’ speech was analyzed for temporal, spectral, and linguistic markers, alongside standardized UX measures, brief mood and stress ratings, and a post-study questionnaire. We found correlations between specific speech features and self-reported satisfaction and experience. Furthermore, a machine learning model trained on speech features achieved promising accuracy in classifying UX levels, indicating that this might be a reasonable alternative to self-report instruments. Our findings establish speech as a viable, real-time signal for implicitly measuring UX and point toward adaptive VUIs that respond dynamically to emotional and usability-related vocal cues. Yong Ma 0003, Xuesong Zhang 0002, Natalia Bartlomiejczyk, Seungwoo Je, Adrian Holzer, Morten Fjeld, Andreas Butz |
CHI | 1 |
| 2026 | Neutral by Default? Replicating User Vocal Responses to Negative Affective Cues in Conversational AgentsabstractConversational agents (CAs) increasingly detect users’ emotions, yet deciding how to respond, especially to negative affect, remains a central design challenge. We conducted a role-switching study in which participants reply as the CAs to simulated users expressing anger, sadness, or fear. Results reveal systematic, gender-linked patterns: most male participants favored a neutral, affect-balanced stance and prioritized clarification or task progress, whereas most female participants produced a wider range of non-neutral responses, more often using explicit empathy, reassurance, and reflective listening. We also observe differences in de-escalation phrasing, validation timing, and follow-up questioning across scenarios. These findings indicate that strategies for handling negative emotions vary with user characteristics and context. Based on these findings, we argue for adaptive CA response policies that calibrate first-turn acknowledgment and information-gathering, tailoring prosody and wording to emotional context in order to support de-escalation, perceived understanding, and user trust. Yong Ma 0003, Yuchong Zhang 0001, Di Fu, Stephanie Zubicueta Portales, Morten Fjeld |
HRI | 1 |
| 2026 | Beyond Information Amount: Rethinking Transparency through Active Reasoning DemandabstractTransparency of a robot's actions and intentions is important for trustful human-robot collaboration. Current approaches for creating transparency through explanations mostly follow the approach "more information creates more transparency", assuming cognitive load remains manageable. However, this ignores human active reasoning, which can, if not too demanding, create a better understanding from less information, as found in education literature. To explore this self-explanation effect, we compared three explanation structures (fully-specified, under-specified, no explanation) that induced different levels of active reasoning demand (low, medium, high). In a controlled laboratory study, 36 participants observed a robot completing rule-based classification tasks of varying difficulty (easy, moderate, hard) across all explanation conditions. We found that a moderate reasoning demand, elicited through under-specified explanations, produced the best understanding of the robot's actions compared to full or no explanations. This challenges the current "more helps more" approach and may help design more effective explanations for creating transparency. Yilu Ye, Yong Ma 0003, Zelun Tony Zhang, Andreas Butz |
HRI | 3 |
| 2026 | "Same Voice, Different Language": An Exploration of Voice-Cloned Translation to Support Non-Native Speakers in Online MeetingsabstractCross-lingual meetings have become essential for global collaboration, yet current translation technologies often strip away vocal identity — the unique speaker characteristics that convey nuance and social presence. While generic text-to-speech (TTS) provides basic intelligibility, it creates a disconnect between speakers and their translated voices, potentially undermining engagement and comprehension. This paper investigates whether voice cloning technology can bridge this gap by preserving speaker identity in real-time translation. We present a controlled study comparing four voice conditions in meeting interpretation: original speech, gender-neutral TTS, gender-matched TTS, and voice cloning. Through a within-subjects experiment with 45 participants, we demonstrate that voice cloning significantly reduces mental workload (p <.001) and enhances user experience across pragmatic quality (p <.001), hedonic quality (p <.001), and overall satisfaction (p <.001) compared to traditional TTS. While original speech maintained advantages in naturalness, voice cloning achieved superior intelligibility, social impression, and user preference. Qualitative analysis revealed that participants valued voice cloning for preserving speaker identity and improving conversation tracking in multi-speaker scenarios. Our findings suggest that identity-preserving translation represents a significant advancement for cross-lingual communication systems, offering both cognitive and social benefits. We conclude with design implications for integrating voice cloning into meeting platforms while addressing ethical considerations around consent and transparency. Yong Ma 0003, Yuchong Zhang 0001, Peter Andrews, Zhikun Wu, Stephanie Zubicueta Portales, Morten Fjeld |
IUI | 1 |
| 2025 | Mind Meets Robots: A Review of EEG-Based Brain-Robot Interaction SystemsabstractBrain-robot interaction (BRI) empowers individuals to control (semi-)automated machines through brain activity, either passively or actively. In the past decade, BRI systems have advanced significantly, primarily leveraging electroencephalogram (EEG) signals. This article presents an up-to-date review of 87 curated studies published between 2018 and 2023, identifying the research landscape of EEG-based BRI systems. The review consolidates methodologies, interaction modes, application contexts, system evaluation, existing challenges, and future directions in this domain. Based on our analysis, we propose a BRI system model comprising three entities: Brain, Robot, and Interaction, depicting their internal relationships. We especially examine interaction modes between human brains and robots, an aspect not yet fully explored. Within this model, we scrutinize and classify current research, extract insights, highlight challenges, and offer recommendations for future studies. Our findings provide a structured design space for human-robot interaction (HRI), informing the development of more efficient BRI frameworks. Yuchong Zhang 0001, Nona Rajabi, Farzaneh Taleb, Andrii Matviienko, Yong Ma 0003, Mårten Björkman, Danica Kragic |
Int. J. Hum. Comput. Interact. | 5 |
| 2024 | Human-centered AI Technologies in Human-robot Interaction for Social SettingsabstractThe increasing integration of human-robot interaction (HRI) into social settings demands the development of human-centered AI technologies that prioritize intuitive, ethical, and empathetic interactions. As robots become more prevalent in everyday life—ranging from assistive devices in healthcare to educational tools in classrooms and customer service agents in retail—it is essential to ensure they can communicate and collaborate with humans in ways that are not only effective but also socially appropriate and meaningful. This workshop aims to explore cutting-edge advancements and interdisciplinary approaches to building AI-driven systems that facilitate effective, meaningful, and socially appropriate interactions between robots and humans across various environments such as healthcare, education, and customer service. We will primarily focus on several key themes, such as human-centered contextual AI, AI-driven intelligent robotics, ethical and responsible AI, and real-world applications. This workshop invites contributions from researchers, practitioners, and developers who are working on AI systems that empower robots to operate effectively in human-centered environments. By addressing challenges such as interpreting human emotions, understanding social cues, and adhering to ethical standards, and by sharing advancements in human-centered AI, we aim to shape the future of HRI. Our goal is to ensure that robots enrich human social experiences, fostering interactions that are not only efficient but also enhance the quality of life. By uniting efforts from various disciplines, we aspire to create robots that seamlessly integrate into society, ultimately contributing to a more harmonious coexistence between humans and robotic systems. Yuchong Zhang 0001, Khaled Kassem, Zhengya Gong, Yong Ma 0003, Emma Kirjavainen, Jonna Häkkilä |
MUM | 5 |
| 2021 | A Journey Through Nature: Exploring Virtual Restorative Environments as a Means to Relax in Confined SpacesabstractVirtual Reality (VR) technologies can counteract stress or fatigue and restore attention, e.g., by recreating the beauty of nature in a Virtual Restorative Environment (VRE). This has gained additional relevance in the current pandemic: When facing the stress of physical restrictions and a limited activity space, how can VR technologies provide the individual experience of being away? We created a VRE that can be used during trips in automated cars using a captured natural environment and simulated artifacts that communicate vehicle information during VR relaxation. In a user study (N = 21), we compared the proposed in-car VRE to simply closing the eyes. We found that the VRE strongly improved the subjective ratings of mood and slightly increased attentional capacity and the objectively measured performance in a working memory test. Our results provide a concrete starting point for exploring calming VR experiences for future passengers, but also users at home. Yong Ma 0003, Puzhen Li, Andreas Butz |
Creativity & Cognition | 2 |
| 2021 | You Sound Relaxed Now - Measuring Restorative Effects from Speech Signals
Yong Ma 0003, Heiko Drewes, Andreas Butz |
INTERACT (2) | 1 |
| 2019 | Automatic Image Segmentation for Microwave Tomography (MWT): From Implementation to Comparative EvaluationabstractInspired by its high performance in image-based medical analysis, this poster paper explores the use of advanced segmentation techniques for industrial Microwave Tomography (MWT). Our context is the visual analysis of moisture levels in porous foams undergoing microwave drying. We propose an automatic segmentation technique---MWT Segmentation based on K-means (MWTS-KM) and demonstrate its efficiency and accuracy for industrial use. MWTS-KM consists of three stages: image augmentation, grey-scale conversion, and K-means implementation. To estimate the performance of this technique, we empirically benchmark its efficiency and accuracy against two well-established alternatives: Otsu and K-means. To elicit performance data, three metrics (Jaccard index, Dice coefficient and false positive) are used. Based on our experiments, our results indicate that MWTS-KM outperforms the well-established Otsu and K-means. Yuchong Zhang 0001, Yong Ma 0003, Adel Omrani, Morten Fjeld, Marco Fratarcangeli |
VINCI | 2 |