Shengyu Hao

dblp:236/6477 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-8652-8556ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
abstract
Understanding dynamic 4D scenes from an egocentric perspective—modeling changes in 3D spatial structure over time—is crucial for human–machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets contain dynamic scenes, they lack unified 4D annotations and task-driven evaluation protocols for fine-grained spatio-temporal reasoning, especially on motion of objects and human, together with their interactions. To address this gap, we introduce EgoDynamic4D, a novel QA benchmark on highly dynamic scenes, comprising RGB-D video, camera poses, globally unique instance masks, and 4D bounding boxes. We construct 927K QA pairs accompanied by explicit Chain-of-Thought (CoT), enabling verifiable, step-by-step spatio-temporal reasoning. We design 12 dynamic QA tasks covering agent motion, human–object interaction, trajectory prediction, relation understanding, and temporal–causal reasoning, with fine-grained, multidimensional metrics. To tackle these tasks, we propose an end-to-end spatio-temporal reasoning framework that unifies dynamic and static scene information, using instance-aware feature encoding, time and camera encoding, and spatially adaptive down-sampling to compress large 4D scenes into token sequences manageable by LLMs. Experiments on EgoDynamic4D show that our method consistently outperforms baselines, validating the effectiveness of multimodal temporal modeling for egocentric dynamic scene understanding.
Shengyu Hao, Bocheng Hu, Hongwei Wang 0001, Gaoang Wang
AAAI2
2026 Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field
Wenhao Hu 0002, Wenhao Chai, Shengyu Hao, Xiaotong Cui, Xuexiang Wen, Jenq-Neng Hwang, Gaoang Wang
Int. J. Comput. Vis.3
2025 RAPID: Recognition of Any-Possible DrIver Distraction via Multi-view Pose Generation Models
abstract
Driver distraction remains a pressing traffic safety issue. Drivers are often careless with their distraction behaviours, which may cause serious traffic accidents. However, current Driver Monitoring Systems (DMS) cannot be put into practical application well, which tend to have high latency, lack precision, and are unable to cover all distraction behaviours. In this paper, we assume driver distraction to be a One-Class Classification (OCC) problem and build an unsupervised learning baseline based on denoising diffusion probabilistic models (DDPM) called RAPID which aggregates future patterns generated by the diffusion process to detect distraction, considering the diversity of normal and abnormal situations. Besides, we propose a skeleton-based synchronized multi-view dataset with diverse distraction behaviours called sktDD (skeleton-based Driver Distraction dataset) to improve on existing datasets. RAPID facilitates a frame-level (0.03 second) and undefined prediction with AUC score beyond State-of-the-Art (SOTA) methods, surpassing currently typical DMS that rely on post-processing procedures and predefined actions. RAPID has the potential to bring significant advancements in the field of traffic safety, which can also be applied in future self-driving scenarios to determine whether the remote-driving operator’s current state is suitable to take over. Our dataset and code are available at https://github.com/jingyulei/rapid.
Jingyu Lei, Shengyu Hao, Gaoang Wang, Der-Horng Lee
ICASSP2
2025 Efficient Transfer From Image-Based Large Multimodal Models to Video Tasks
abstract
Extending image-based Large Multimodal Models (LMMs) to video-based LMMs always requires temporal modeling in the pre-training. However, training the temporal modules gradually erases the knowledge of visual features learned from various image-text-based scenarios, leading to degradation in some downstream tasks. % Adapting pre-trained video-based large language models (LLMs) to downstream fine-grained video understanding tasks always requires modeling on temporal modules. However, training the temporal modules during video pretraining gradually erases the knowledge of visual features learned from various image-text-based scenarios, leading to degradation in some downstream tasks. % Instead of tuning video-based LLMs to downstream tasks, To address this issue, in this paper, we introduce a novel, efficient transfer approach termed MTransLLAMA, which employs transfer learning from pre-trained image LMMs for fine-grained video tasks with only small-scale training sets. Our method enablesfewer trainable parametersand achievesfaster adaptationandhigher accuracythan pre-training video-based LMM models. Specifically, our method adopts early fusion between textual and visual features to capture fine-grained information, reuses spatial attention weights in temporal attentions for cyclical spatial-temporal reasoning, and introduces dynamic attention routing to capture both global and local information in spatial-temporal attentions. Experiments demonstrate that across multiple datasets and tasks, without relying on video pre-training, our model achieves state-of-the-art performance, enabling lightweight and efficient transfer from image-based LMMs to fine-grained video tasks.
Shidong Cao, Zhonghan Zhao, Shengyu Hao, Wenhao Chai, Jenq-Neng Hwang, Hongwei Wang 0001, Gaoang Wang
IEEE Trans. Multim.3
2025 A Survey of Deep Learning in Sports Applications: Perception, Comprehension, and Decision
abstract
Deep learning has the potential to revolutionize sports performance, with applications ranging from perception and comprehension to decision. This article presents a comprehensive survey of deep learning in sports performance, focusing on three main aspects: algorithms, datasets and virtual environments, and challenges. First, we discuss the hierarchical structure of deep learning algorithms in sports performance which includes perception, comprehension and decision while comparing their strengths and weaknesses. Second, we list widely used existing datasets in sports and highlight their characteristics and limitations. Finally, we summarize current challenges and point out future trends of deep learning in sports. Our survey provides valuable reference material for researchers interested in deep learning in sports applications.
Zhonghan Zhao, Wenhao Chai, Shengyu Hao, Wenhao Hu 0002, Guanhong Wang, Shidong Cao, Mingli Song, Jenq-Neng Hwang, Gaoang Wang
IEEE Trans. Vis. Comput. Graph.3
2024 See and Think: Embodied Agent in Virtual Environment
Zhonghan Zhao, Wenhao Chai, Boyi Li 0002, Shengyu Hao, Shidong Cao, Tian Ye 0001, Gaoang Wang
ECCV (8)5
2024 Vision meets mmWave Radar: 3D Object Perception Benchmark for Autonomous Driving
abstract
Sensor fusion is crucial for an accurate and robust perception system on autonomous vehicles. Most existing datasets and perception solutions focus on fusing cameras and LiDAR. However, the collaboration between camera and radar is significantly under-exploited. Incorporating rich semantic information from the camera and reliable 3D information from the radar can achieve an efficient, cheap, and portable solution for 3D perception tasks. It can also be robust to different lighting or all-weather driving scenarios due to the capability of mmWave radars. In this paper, we introduce the CRUW3D dataset, including 66K synchronized and well-calibrated camera, radar, and LiDAR frames in various driving scenarios. Unlike other large-scale autonomous driving datasets, our radar data is in the format of radio frequency (RF) tensors that contain not only 3D location information but also spatio-temporal semantic information. This kind of radar format can enable machine learning models to generate more reliable object perception results after interacting and fusing the information or features between the camera and radar. We run several camera- and radar-based baseline methods for 3D object detection and multi-object tracking on our dataset. We hope the CRUW3D dataset will foster radar and multi-modal 3D perception research. CRUW3D is available at https://huggingface.co/datasets/uwipl/CRUW3D
Yizhou Wang 0005, Jen-Hao Cheng, Jui-Te Huang, Sheng-Yao Kuan, Qiqian Fu, Chiming Ni, Shengyu Hao, Gaoang Wang, Guanbin Xing, Hui Liu 0011, Jenq-Neng Hwang
IV7
2024 Ego3DT: Tracking Every 3D Object in Ego-centric Videos
abstract
The growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localization and tracking of objects in ego-centric videos, primarily due to the substantial variability in viewing angles. Addressing this issue, this paper introduces a novel zero-shot approach for the 3D reconstruction and tracking of all objects from the ego-centric video. We present Ego3DT, a novel framework that initially identifies and extracts detection and segmentation information of objects within the ego environment. Utilizing information from adjacent video frames, Ego3DT dynamically constructs a 3D scene of the ego view using a pre-trained 3D scene reconstruction model. Additionally, we have innovated a dynamic hierarchical association mechanism for creating stable 3D tracking trajectories of objects in ego-centric videos. Moreover, the efficacy of our approach is corroborated by extensive experiments on two newly compiled datasets, with 1.04 × - 2.90× in HOTA, showcasing the robustness and accuracy of our method in diverse ego-centric scenarios.
Shengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun, Wendi Hu, Jieyang Zhou, Yixian Zhao, Yizhou Wang 0005, Gaoang Wang
ACM Multimedia1
2024 DIVOTrack: A Novel Dataset and Baseline Method for Cross-View Multi-Object Tracking in DIVerse Open Scenes
Shengyu Hao, Peiyuan Liu, Yibing Zhan, Kaixun Jin, Zuozhu Liu, Mingli Song, Jenq-Neng Hwang, Gaoang Wang
Int. J. Comput. Vis.1
2024 DiffFashion: Reference-Based Fashion Design With Structure-Aware Transfer by Diffusion Models
abstract
Image-based fashion design with AI techniques has attracted increasing attention in recent years. We focus on the reference-based fashion design task, where we aim to combine a reference appearance image and a clothing image to generate a new fashion clothing image. Although existing diffusion-based image translation methods have enabled flexible style transfer, it is often difficult to transfer the appearance of the image realistically during reverse diffusion. When the referenced appearance domain greatly differs from the source domain, it often leads to the collapse in the translation. To tackle this issue, we present a novel diffusion model-based unsupervised structure-aware transfer method, namelyDiffFashion. Our method is free of model tuning and structure-preserving and has high flexibility in transferring from images with large domain gaps. Specifically, based on the optimal transport properties, we keep a shared latent across the clothing image and reference appearance image to bridge the gap between the two domains in the denoising process, and the latent of the reference image is gradually adapted to the clothing domain. Simultaneously, the structure is transferred from the source clothing to the output fashion image with mixed guidance, including pre-trained Vision Transformer (ViT) guidance and a foreground mask guidance, to further preserve the structure and appearance semantics from source and reference images. Our experimental results show that the proposed method outperforms state-of-the-art baseline models, generating more realistic images in the fashion design task.
Shidong Cao, Wenhao Chai, Shengyu Hao, Yanting Zhang 0001, Hangyue Chen, Gaoang Wang
IEEE Trans. Multim.3
2021 Weakly supervised instance segmentation using multi-prior fusion
Shengyu Hao, Gaoang Wang, Renshu Gu
Comput. Vis. Image Underst.1