Jouh Yeong Chew

dblp:53/7956 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-7906-2113ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Beyond Reciprocity: Psychological Needs as a Foundation for Human-AI Cooperation
abstract
In the coming years, human-AI cooperation will become an even more central part of many people’s private and work lives, supporting and shaping cognitive processes, while also influencing social dynamics. To understand the conditions under which humans choose to cooperate with AI, researchers have widely utilized evolutionary theories and game-theoretic approaches. However, these frameworks primarily emphasize utility maximization and strategic behavior, overlooking the subjective, experiential dimension of cooperation. To overcome this limitation, we here propose an integrated framework for the emergence of human-AI cooperation, which combines a mechanistic layer drawn from evolutionary theories of cooperation with an experiential layer provided by self-determination theory. We further propose to operationalize the human cooperation intent as the perceived balance between benefits and costs of cooperating, which is moderated by the extent to which psychological needs are satisfied through the cooperation. Our framework offers a novel approach for experimentally testing the formation a human-AI cooperation intent, highlighting not only when and why cooperation may occur, but also how it can be designed to be intrinsically motivating and meaningful for users.
Christiane Attig, Alan Sarkisian, Jouh Yeong Chew, Christiane B. Wiebel-Herboth
HAI3
2025 Workshop on Socially Aware and Cooperative Intelligent Systems
abstract
This workshop theme centers on the development of AI agents and systems that are capable of understanding, adapting to, and reacting to collaborate with humans in compliance with the social norms. These systems leverage insights from social psychology, cognitive science, robotics, and AI to interpret social cues, anticipate the needs of others, and coordinate actions effectively within dynamic and often unpredictable contexts. We focus on embedding social awareness into AI systems, leading to Cooperative Intelligence [23] which focuses on building trust and relationship between humans and intelligent systems, instead of focusing on functions to replace humans. This paradigm is expected to realize a hybrid society, where humans coexist with ubiquitous intelligent agents.
Jouh Yeong Chew, Alan Sarkisian, Christiane B. Wiebel-Herboth, Christiane Attig, Zhaobo Zheng, Shigeaki Nishina
HAI1
2025 SpeechCAT: Cross-Attentive Transformer for Audio to Motion Generation
abstract
Audio-to-motion generation is an important task with applications in virtual avatar creation for XR systems and intelligent robot control in daily life scenarios. However, most existing motion generation methods rely on a single encoder-decoder architecture to model all body parts simultaneously, which limits their ability to capture the diverse and complex motions exhibited by humans. In this paper, we propose a novel method, SpeechCAT, that employs three separate encoder-decoder modules to individually model the motions of the face, body, and hands. To capture the relationships and synchronization among these body parts, we introduce a cross-attention mechanism to effectively learn their correlations. SpeechCAT ensures sufficient capacity to model the unique characteristics of each body part while preserving the coherence between them. Our experimental results demonstrate the superiority of SpeechCAT over baseline methods, highlighting its effectiveness in generating diverse, realistic, and synchronized motions with face, body, and hand parts.
Sebastian Deaconu, Xiangwei Shi, Thomas Markhorst, Jouh Yeong Chew, Xucong Zhang
HRI4
2025 Learning Nonverbal Cues in Multiparty Social Interactions for Robotic Facilitators
abstract
Conventional behavior cloning (BC) models often struggle to replicate the subtleties of human actions. Previous studies have attempted to address this issue through the development of a new BC technique: Implicit Behavior Cloning (IBC). This new technique consistently outperformed the conventional Mean Squared Error (MSE) BC models in a variety of tasks. Our goal is to replicate the performance of the IBC model by Florence [in Proceedings of the 5th Conference on Robot Learning, 164:158–168, 2022], for social interaction tasks using our custom dataset. While previous studies have explored the use of large language models (LLMs) for enhancing group conversations, they often overlook the significance of non-verbal cues, which constitute a substantial part of human communication. We propose using IBC to replicate nonverbal cues like gaze behaviors. The model is evaluated against various types of facilitator data and compared to an explicit, MSE BC model. Results show that the IBC model outperforms the MSE BC model across session types using the same metrics used in the previous IBC paper. Despite some metrics showing mixed results which are explainable for the custom dataset for social interaction, we successfully replicated the IBC model to generate nonverbal cues. Our contributions are (1) the replication and extension of the IBC model, and (2) a nonverbal cues generation model for social interaction. These advancements facilitate the integration of robots into the complex interactions between robots and humans, e.g., in the absence of a human facilitator.
Antonio Lech Martin-Ozimek, Isuru Jayarathne, Su Larb Mon, Jouh Yeong Chew
HRI4
2025 Diffusion-Based Imitation Learning for Social Pose Generation
abstract
Intelligent agents, such as robots and virtual agents, must understand the dynamics of complex social interactions to interact with humans. Effectively representing social dynamics is challenging because we require multi-modal, synchronized observations to understand a scene. We explore how using a single modality, the pose behavior, of multiple individuals in a social interaction can be used to generate nonverbal social cues for the facilitator of that interaction. The facilitator acts to make a social interaction proceed smoothly and is an essential role for intelligent agents to replicate in human-robot interactions. In this paper, we adapt an existing diffusion behavior cloning model to learn and replicate facilitator behaviors. Furthermore, we evaluate two representations of pose observations from a scene, one representation has pre-processing applied and one does not. The purpose of this paper is to introduce a new use for diffusion behavior cloning for pose generation in social interactions. The second is to understand the relationship between performance and computational load for generating social pose behavior using two different techniques for collecting scene observations. As such, we are essentially testing the effectiveness of two different types of conditioning for a diffusion model. We then evaluate the resulting generated behavior from each technique using quantitative measures such as mean per-joint position error (MPJPE), training time, and inference time. Additionally, we plot training and inference time against MPJPE to examine the trade-offs between efficiency and performance. Our results suggest that the further pre-processed data can successfully condition diffusion models to generate realistic social behavior, with reasonable trade-offs in accuracy and processing time. Future work will focus on extending this approach to generate multiple nonverbal social cues, on generalizing this method to multiple types of social activity, and on evaluating the results using human evaluators using a virtual agent or robot as a facilitator.
Antonio Lech Martin-Ozimek, Isuru Jayarathne, Su Larb Mon, Jouh Yeong Chew
HRI4
2025 GazeHTA: End-to-End Gaze Target Detection with Head-Target Association
Zhiyi Lin 0003, Jouh Yeong Chew, Jan C. van Gemert, Xucong Zhang
ICRA2
2024 Modeling social interaction dynamics using temporal graph networks
abstract
Integrating intelligent systems, such as robots, into dynamic group settings poses challenges due to the mutual influence of human behaviors and internal states. A robust representation of social interaction dynamics is essential for effective human-robot collaboration. Existing approaches often narrow their focus to facial expressions or speech, overlooking the broader context. We propose employing an adapted Temporal Graph Networks to comprehensively represent social interaction dynamics while enabling its practical implementation. Our method incorporates temporal multi-modal behavioral data including gaze interaction, voice activity and environmental context. This representation of social interaction dynamics is trained as a link prediction problem using annotated gaze interaction data. The F1-score outperformed the baseline model by 37.0%. This improvement is consistent for a secondary task of next speaker prediction which achieves an improvement of 29.0%. Our contributions are two-fold, including a model to representing social interaction dynamics which can be used for many downstream human-robot interaction tasks like human state inference and next speaker prediction. More importantly, this is achieved using a more concise yet efficient message-passing method, significantly reducing the message size from 768 to 14 while outperforming the baseline model.
Joanne Taery Kim, Archit Naik, Isuru Jayarathne, Sehoon Ha, Jouh Yeong Chew
RO-MAN5