Shaohua Pan 0002

dblp:22/7025-2 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-6261-5268ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 68% Image and video processing · 18% Virtual and augmented reality · 14%
Artificial intelligence
2 papers
3D vision · 64% Face, body and person analysis · 36%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › pose estimation › generative pose estimation
diffusion-based pose estimation
0.912025
DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera · IEEE Trans. Vis. Comput. Graph. 2025
Computer vision › Face, body and person analysis
human pose estimation
0.912025
DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera · IEEE Trans. Vis. Comput. Graph. 2025
Computer animation and physical simulation › motion capture
human motion capture
0.912025
DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera · IEEE Trans. Vis. Comput. Graph. 2025
Image and video processing › motion analysis › human motion analysis
human motion estimation
0.912025
Improving Global Motion Estimation in Sparse IMU-based Motion Capture with Physics · ACM Trans. Graph. 2025
Computer animation and physical simulation
motion capture
0.912025
DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera · IEEE Trans. Vis. Comput. Graph. 2025
Computer vision › 3D vision › motion capture
human motion capture
0.712023
Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture · SIGGRAPH Asia 2023
Computer animation and physical simulation › motion capture
inertial motion capture
0.712023
EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted Sensors · ACM Trans. Graph. 2023
Virtual and augmented reality › tracking and registration
simultaneous localization and mapping
0.712023
EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted Sensors · ACM Trans. Graph. 2023
Wearable and physiological sensing › inertial sensing
inertial measurement unit
0.712023
Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture · SIGGRAPH Asia 2023
Wearable and physiological sensing
motion capture
0.712023
Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture · SIGGRAPH Asia 2023

Methods — techniques the papers use, named apart from their topics

monocular camera · 1.7diffusion model · 1.7IMU fusion · 1.7hidden state feedback · 1.3dual coordinate strategy · 1.3physical optimization · 0.9deep learning · 0.9inertial mocap fusion · 0.7SLAM · 0.7
YearPublicationVenuePosition
2025 Improving Global Motion Estimation in Sparse IMU-based Motion Capture with Physics
abstract
By learning human motion priors, motion capture can be achieved by 6 inertial measurement units (IMUs) in recent years with the development of deep learning techniques, even though the sensor inputs are sparse and noisy. However, human global motions are still challenging to be reconstructed by IMUs. This paper aims to solve this problem by involving physics. It proposes a physical optimization scheme based on multiple contacts to enable physically plausible translation estimation in the full 3D space where the z-directional motion is usually challenging for previous works. It also considers gravity in local pose estimation which well constrains human global orientations and refines local pose estimation in a joint estimation manner. Experiments demonstrate that our method achieves more accurate motion capture for both local poses and global motions. Furthermore, by deeply integrating physics, we can also estimate 3D contact, contact forces, joint torques, and interacting proxy surfaces. Code is available at https://xinyu-yi.github.io/GlobalPose/.
Xinyu Yi, Shaohua Pan 0002, Feng Xu 0005
ACM Trans. Graph.2
2025 DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera
abstract
Combining sparse IMUs and a monocular camera is a new promising setting to perform real-time human motion capture. This paper proposes a diffusion-based solution to learn human motion priors and fuse the two modalities of signals together seamlessly in a unified framework. By delicately considering the characteristics of the two signals, the sequential visual information is considered as a whole and transformed into a condition embedding, while the inertial measurement is concatenated with the noisy body pose frame by frame to construct a sequential input for the diffusion model. Firstly, we observe that the visual information may be unavailable in some frames due to occlusions or subjects moving out of the camera view. Thus incorporating the sequential visual features as a whole to get a single feature embedding is robust to the occasional degenerations of visual information in those frames. On the other hand, the IMU measurements are robust to occlusions and always stable when signal transmission has no problem. So incorporating them frame-wisely could better explore the temporal information for the system. Experiments have demonstrated the effectiveness of the system design and its state-of-the-art performance in pose estimation compared with the previous works. The code will be released.
Shaohua Pan 0002, Xinyu Yi, Yan Zhou 0003, Weihua Jian, Yuan Zhang 0020, Pengfei Wan 0001, Feng Xu 0005
IEEE Trans. Vis. Comput. Graph.1
2023 Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture
abstract
Either RGB images or inertial signals have been used for the task of motion capture (mocap), but combining them together is a new and interesting topic. We believe that the combination is complementary and able to solve the inherent difficulties of using one modality input, including occlusions, extreme lighting/texture, and out-of-view for visual mocap and global drifts for inertial mocap. To this end, we propose a method that fuses monocular images and sparse IMUs for real-time human motion capture. Our method contains a dual coordinate strategy to fully explore the IMU signals with different goals in motion capture. To be specific, besides one branch transforming the IMU signals to the camera coordinate system to combine with the image information, there is another branch to learn from the IMU signals in the body root coordinate system to better estimate body poses. Furthermore, a hidden state feedback mechanism is proposed for both two branches to compensate for their own drawbacks in extreme input cases. Thus our method can easily switch between the two kinds of signals or combine them in different cases to achieve a robust mocap. Quantitative and qualitative results demonstrate that by delicately designing the fusion method, our technique significantly outperforms the state-of-the-art vision, IMU, and combined methods on both global orientation and local pose estimation. Our codes are available for research at https://shaohua-pan.github.io/robustcap-page/.
Shaohua Pan 0002, Xinyu Yi, Xingkang Zhou, Jijunnan Li, Feng Xu 0005
SIGGRAPH Asia1
2023 EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted Sensors
abstract
Human and environment sensing are two important topics in Computer Vision and Graphics. Human motion is often captured by inertial sensors, while the environment is mostly reconstructed using cameras. We integrate the two techniques together in EgoLocate, a system that simultaneously performs human motion capture (mocap), localization, and mapping in real time from sparse body-mounted sensors, including 6 inertial measurement units (IMUs) and a monocular phone camera. On one hand, inertial mocap suffers from large translation drift due to the lack of the global positioning signal. EgoLo-cate leverages image-based simultaneous localization and mapping (SLAM) techniquesto locate the human in the reconstructed scene. Onthe other hand, SLAM often fails when the visual feature is poor. EgoLocate involves inertial mocap to provide a strong prior for the camera motion. Experiments show that localization, a key challenge for both two fields, is largely improved by our technique, compared with the state of the art of the two fields. Our codes are available for research at https://xinyu-yi.github.io/EgoLocate/.
Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Vladislav Golyanik, Shaohua Pan 0002, Christian Theobalt, Feng Xu 0005
ACM Trans. Graph.5