Shaun Jing Heng Ong

dblp:429/6771 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0005-7430-8467ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 44% Haptics and multimodal interaction · 44% Wearable and physiological sensing · 13%
Artificial intelligence
1 paper
Face, body and person analysis · 100%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation
1.012026
EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality · IEEE Trans. Vis. Comput. Graph. 2026
Computer vision › Face, body and person analysis
human pose estimation
1.012026
EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality · IEEE Trans. Vis. Comput. Graph. 2026
Virtual and augmented reality › tracking
full-body tracking
1.012026
EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality · IEEE Trans. Vis. Comput. Graph. 2026
Virtual and augmented reality
immersive interaction
1.012026
EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality · IEEE Trans. Vis. Comput. Graph. 2026
Haptics and multimodal interaction › multimodal communication
multimodal expression
1.012026
Signals of Aggression: Modelling Multimodal Cues and Perceptual Effects in Virtual Agents · CHI 2026
Human-AI interaction
virtual agents
1.012026
Signals of Aggression: Modelling Multimodal Cues and Perceptual Effects in Virtual Agents · CHI 2026

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 3.0kinematic optimization · 3.0cross-attention · 3.0user study · 1.0
YearPublicationVenuePosition
2026 Signals of Aggression: Modelling Multimodal Cues and Perceptual Effects in Virtual Agents
abstract
Aggression is a socially complex behaviour that intelligent virtual agents (IVAs) must convincingly convey in applications such as customer service and conflict training. Despite its importance, aggression remains understudied: prior work has focused on basic emotions and unimodal cues, providing little insight into how aggression can be modelled multimodally or systematically scaled by intensity. We present a psychologically grounded model that parametrises language, voice, body movement and facial expressions, across four aggression levels. We evaluated the model in two studies with 38 flight attendants. Experiment 1 tested unimodal cues, showing all modalities except language conveyed aggression gradients. Experiment 2 extended this by combining modalities, demonstrating that coordinated multimodal integration stabilised weaker language cues and produced perceptually robust aggression levels (low, mid, and high) with body and facial cues carrying most weight. Our work contributes the first validated multimodal, multi-level aggression model for IVAs, offering design principles for broader socially expressive agents.
Shaun Jing Heng Ong, Aiden Koh, Shaoyu Cai, Felicia Fang-Yi Tan, Patrick Chia, Eng Tat Khoo
CHI1
2026 EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality
abstract
Immersive virtual reality (VR) applications demand accurate, temporally coherent full-body pose tracking. Recent head-mounted camera-based approaches show promise in egocentric pose estimation, but encounter challenges when applied to VR head-mounted displays (HMDs), including temporal instability, inaccurate lower-body estimation, and the lack of real-time inference. To address these limitations, we present EgoPoseVR, an end-to-end framework for accurate egocentric full-body pose estimation in VR that integrates headset motion cues with egocentric RGB-D observations through a dual-modality fusion pipeline. A spatiotemporal encoder extracts frame- and joint-level representations, which are fused via cross-attention to fully exploit complementary motion cues across modalities. A kinematic optimization module then imposes constraints from HMD signals, enhancing the accuracy and stability of pose estimation. To facilitate training and evaluation, we introduce a large-scale synthetic dataset of over 1.8 million temporally aligned HMD and RGB-D frames across diverse VR scenarios. Experimental results show that EgoPoseVR outperforms state-of-the-art egocentric pose estimation models. A user study in real-world scenes further shows that EgoPoseVR achieved significantly higher subjective ratings in accuracy, stability, embodiment, and intention for future use compared to baseline methods. These results show that EgoPoseVR enables robust full-body pose tracking, offering a practical solution for accurate VR embodiment without requiring additional body-worn sensors or room-scale tracking systems.
Haojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Koh, Fuxi Ouyang, Eng Tat Khoo
IEEE Trans. Vis. Comput. Graph.2