Chuxuan Zhang

dblp:219/5082 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-8993-5694ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 A novel graph convolutional network with multi-kernel learning approach for predicting snoRNA-disease associations
Chuxuan Zhang, Yuhan Fan, Caiyi Liang, Sibo Wen, Ruihan Lai
Expert Syst. Appl.1
2024 Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
abstract
The emotional theory of mind problem requires facial expressions, body pose, contextual information and implicit commonsense knowledge to reason about the person's emotion and its causes, making it currently one of the most difficult problems in affective computing. In this work, we propose multiple methods to incorporate the emotional reasoning capabilities by constructing “narrative captions” relevant to emotion perception, that includes contextual and physical signal descriptors that focuses on “Who”, “What”, “Where” and “How” questions related to the image and emotions of the individual. We propose two distinct ways to construct these captions using zero-shot classifiers (CLIP) and fine-tuning visual-language models (LLaVA) over human generated descriptors. We further utilize these captions to guide the reasoning of language (GPT-4) and vision-language models (LLa Va, GPT-Vision). We evaluate the use of the resulting models in an image-to-language-to-emotion task. Our experiments showed that combining the “Fast” narrative descriptors and “Slow” reasoning of language models is a promising way to achieve emotional theory of mind.
Yasaman Etesam, Özge Nilay Yalçin, Chuxuan Zhang, Angelica Lim
ACII3
2024 Contextual Emotion Recognition using Large Vision Language Models
abstract
How does the person in the bounding box feel?" Achieving human-level recognition of the apparent emotion of a person in real world situations remains an unsolved task in computer vision. Facial expressions are not enough: body pose, contextual knowledge, and commonsense reasoning all contribute to how humans perform this emotional theory of mind task. In this paper, we examine two major approaches enabled by recent large vision language models: 1) image captioning followed by a language-only LLM, and 2) vision language models, under zero-shot and fine-tuned setups. We evaluate the methods on the Emotions in Context (EMOTIC) dataset and demonstrate that a vision language model, fine-tuned even on a small dataset, can significantly outperform traditional baselines. The results of this work aim to help robots and agents perform emotionally sensitive decision-making and interaction in the future.
Yasaman Etesam, Özge Nilay Yalçin, Chuxuan Zhang, Angelica Lim
IROS3
2024 React to This! How Humans Challenge Interactive Agents using Nonverbal Behaviors
abstract
How do people use their faces and bodies to test the interactive abilities of a robot? Making lively, believable agents is often seen as a goal for robots and virtual agents but believability can easily break down. In this Wizard-of-Oz (WoZ) study, we observed 1169 nonverbal interactions between 20 participants and 6 types of agents. We collected the nonverbal behaviors participants used to challenge the characters physically, emotionally, and socially. The participants interacted freely with humanoid and non-humanoid forms: a robot, a human, a penguin, a pufferfish, a banana, and a toilet. We present a human behavior codebook of 188 unique nonverbal behaviors used by humans to test the virtual characters. The insights and design strategies drawn from video observations aim to help build more interaction-aware and believable robots, especially when humans push them to their limits.
Chuxuan Zhang, Bermet Burkanova, Lawrence H. Kim, Lauren Yip, Ugo Cupcic, Stéphane Lallée, Angelica Lim
IROS1
2023 Contextual Emotion Estimation from Image Captions
abstract
Emotion estimation in images is a challenging task, typically using computer vision methods to directly estimate people’s emotions using face, body pose and contextual cues. In this paper, we explore whether Large Language Models (LLMs) can support the contextual emotion estimation task, by first captioning images, then using an LLM for inference. First, we must understand: how well do LLMs perceive human emotions? And which parts of the information enable them to determine emotions? One initial challenge is to construct a caption that describes a person within a scene with information relevant for emotion perception. Towards this goal, we propose a set of natural language descriptors for faces, bodies, interactions, and environments. We use them to manually generate captions and emotion annotations for a subset of 331 images from the EMOTIC dataset. These captions offer an interpretable representation for emotion estimation, towards understanding how elements of a scene affect emotion perception in LLMs and beyond. Secondly, we test the capability of a large language model to infer an emotion from the resulting image captions. We find that GPT3.5, specifically the text-davinci-003 model, provides surprisingly reasonable emotion predictions consistent with human annotations, but accuracy can depend on the emotion concept. Overall, the results suggest promise in the image captioning and LLM approach.
Vera Yang, Archita Srivastava, Yasaman Etesam, Chuxuan Zhang, Angelica Lim
ACII4
2023 Read the Room: Adapting a Robot's Voice to Ambient and Social Contexts
abstract
How should a robot speak in a formal, quiet and dark, or a bright, lively and noisy environment? By designing robots to speak in a more social and ambient-appropriate manner we can improve perceived awareness and intelligence for these agents. We describe a process and results toward selecting robot voice styles for perceived social appropriateness and ambiance awareness. Understanding how humans adapt their voices in different acoustic settings can be challenging due to difficulties in voice capture in the wild. Our approach includes 3 steps: (a) Collecting and validating voice data interactions in virtual Zoom ambiances, (b) Exploration and clustering human vocal utterances to identify primary voice styles, and (c) Testing robot voice styles in recreated ambiances using projections, lighting and sound. We focus on food service scenarios as a proof-of-concept setting. We provide results using the Pepper robot's voice with different styles, towards robots that speak in a contextually appropriate and adaptive manner. Our results with N=120 participants provide evidence that the choice of voice style in different ambiances impacted a robot's perceived intelligence in several factors including: social appropriateness, comfort, awareness, human-likeness and competency.
Paige Tuttosi, Emma Hughson, Akihiro Matsufuji, Chuxuan Zhang, Angelica Lim
IROS4
2022 Choose or Fuse: Enriching Data Views with Multi-label Emotion Dynamics
abstract
Many emotion classification and prediction approaches focus on emotion state, defined as static and single-valued. In contrast, our in-body experience is of sensations that can quickly evolve, consistent with scientific evidence of physiological regulation mechanisms. Can we reframe classification to estimate dynamic emotion parameters at interactive rates? For insight into dynamic emotion characteristics, we developed a multipass labelling protocol to capture controlled yet genuine emotion evolution elicited as 16 participants played a tense video game. We analyze and align multiple self-report outputs, inspect the signals for emotion dynamics, and consider label metaphors of position and angle — “where I am” vs. “where I'm going”. Finally, we reflect on the benefits and drawbacks of such a protocol for developing models of fast-evolving emotion.
Laura Cang, Rúbia Reis Guerra, Paul Bucci, Bereket Guta, Karon E. MacLean, Laura Rodgers, Hailey Mah, Shinmin Hsu, Qianqian Feng, Chuxuan Zhang, Anushka Agrawal
ACII10
2021 A Privacy-sensitive Service Selection Method Based on Artificial Fish Swarm Algorithm in the Internet of Things
Bing Jia, Lifei Hao, Chuxuan Zhang, Baoqi Huang
Mob. Networks Appl.3
2019 An IoT Service Aggregation Method Based on Dynamic Planning for QoE Restraints
Bing Jia, Lifei Hao, Chuxuan Zhang, Huili Zhao, Khan Muhammad 0001
Mob. Networks Appl.3
2019 Correction to: An IoT Service Aggregation Method Based on Dynamic Planning for QoE Restraints
Bing Jia, Lifei Hao, Chuxuan Zhang, Huili Zhao, Khan Muhammad 0001
Mob. Networks Appl.3