Zhen Fan 0015

dblp:372/0124 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 41% Face, body and person analysis · 27% Generative modeling · 24%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%
Computer graphics and multimedia
2 papers
Virtual and augmented reality · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
human pose estimation
1.622025
EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs · AAAI 2025
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations · CVPR 2024
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation
0.912025
EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs · AAAI 2025
Wearable and physiological sensing › inertial sensing
inertial measurement unit
0.912025
EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs · AAAI 2025
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.812024
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors · NeurIPS 2024
Computer vision › 3D vision
3d human reconstruction
0.812024
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors · NeurIPS 2024
Computer vision › 3D vision › motion capture
full-body motion capture
0.812024
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations · CVPR 2024
Computer vision › 3D vision › human mesh recovery
human body shape estimation
0.812024
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations · CVPR 2024
Computer vision › Video understanding and tracking › object tracking
human motion tracking
0.812024
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations · CVPR 2024
Machine learning › Generative modeling › diffusion model › 3d-aware diffusion
multi-view diffusion
0.812024
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors · NeurIPS 2024
Machine learning › Generative modeling › image generation
person image synthesis
0.812024
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors · NeurIPS 2024
Computer vision › 3D vision › 3d human reconstruction
single-view human reconstruction
0.812024
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors · NeurIPS 2024
Virtual and augmented reality › immersive interaction › VR interaction
head-mounted display interaction
0.212024
HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations · CVPR 2024

Methods — techniques the papers use, named apart from their topics

temporal feature encoding · 2.6multimodal fusion · 2.6MLP regression · 2.6temporal-spatial feature learning · 1.5sparse IMU fusion · 1.5structure prior · 0.8latent reconstruction transformer · 0.83d gaussian splatting · 0.8
YearPublicationVenuePosition
2025 EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs
abstract
Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to self-occlusion in images or the sparseness and drift of inertial sensors. Most importantly, the lack of real-world datasets containing both modalities is a major obstacle to progress in this field. To overcome the barrier, we propose EMHI, a multimodal Egocentric human Motion dataset with Head-Mounted Display (HMD) and body-worn IMUs, with all data collected under the real VR product suite. Specifically, EMHI provides synchronized stereo images from downward-sloping cameras on the headset and IMU data from body-worn sensors, along with pose annotations in SMPL format. This dataset consists of 885 sequences captured by 58 subjects performing 39 actions, totaling about 28.5 hours of recording. We evaluate the annotations by comparing them with optical marker-based SMPL fitting results. To substantiate the reliability of our dataset, we introduce MEPoser, a new baseline method for multimodal egocentric HPE, which employs a multimodal fusion encoder, temporal feature encoder, and MLP-based regression heads. The experiments on EMHI show that MEPoser outperforms existing single-modal methods and demonstrates the value of our dataset in solving the problem of egocentric HPE. We believe the release of EMHI and the method could advance the research of egocentric HPE and expedite the practical implementation of this technology in VR/AR products.
Zhen Fan 0015, Zhuo Su 0006, Jiarui Zhang 0007, Tianyuan Du, Guidong Wang
AAAI1
2024 HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations
abstract
It is especially challenging to achieve real-time human motion tracking on a standalone VR Head-Mounted Dis-play (HMD) such as Meta Quest and PICO. In this pa-per, we propose HMD-Poser, the first unified approach to recover full-body motions using scalable sparse observations from HMD and body-worn 1MUs. In particular, it can support a variety of input scenarios, such as HMD, HMD+2IMUs, HMD+3IMUs, etc. The scalability of in-puts may accommodate users' choices for both high tracking accuracy and easy-to-wear. A lightweight temporal-spatial feature learning network is proposed in HMD-Poser to guarantee that the model runs in real-time on HMDs. Furthermore, HMD-Poser presents online body shape es-timation to improve the position accuracy of body joints. Extensive experimental results on the challenging AMASS dataset show that HMD-Poser achieves new state-of-the-art results in both accuracy and real-time performance. We also build a new free-dancing motion dataset to evaluate HMD-Poser's on-device performance and investigate the performance gap between synthetic data and real-captured sensor data. Finally, we demonstrate our HMD-Poser with a real-time Avatar-driving application on a commercial HMD. Our code and free-dancing motion dataset are available here.
Zhen Fan 0015, Tianyuan Du, Zhuo Su 0006, Xiaozheng Zheng
CVPR4
2024 HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors
abstract
Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader scenarios. To tackle these issues, we present **HumanSplat**, which predicts the 3D Gaussian Splatting properties of any human from a single input image in a generalizable manner. Specifically, HumanSplat comprises a 2D multi-view diffusion model and a latent reconstruction Transformer with human structure priors that adeptly integrate geometric priors and semantic features within a unified framework. A hierarchical loss that incorporates human semantic information is devised to achieve high-fidelity texture modeling and impose stronger constraints on the estimated multiple views. Comprehensive experiments on standard benchmarks and in-the-wild images demonstrate that HumanSplat surpasses existing state-of-the-art methods in achieving photorealistic novel-view synthesis. Project page: https://humansplat.github.io.
Panwang Pan, Zhuo Su 0006, Chenguo Lin, Zhen Fan 0015, Tingting Shen, Yadong Mu, Yebin Liu
NeurIPS4