Yudi Dai

dblp:229/5710 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-0546-523XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World Coordinate
abstract
Human Motion Recovery (HMR) research mainly focuses on ground-based motions such as running. The study on capturing climbing motion, an off-ground motion, is sparse. This is partly due to the limited availability of climbing motion datasets, especially large-scale and challenging 3D labeled datasets. To address the insufficiency of climbing motion datasets, we collect AscendMotion, a large-scale well-annotated, and challenging climbing motion dataset. It consists of 412k RGB, LiDAR frames, and IMU measurements, including the challenging climbing motions of 22 skilled climbing coaches across 12 different rock walls. Capturing the climbing motions is challenging as it requires precise recovery of not only the complex pose but also the global position of climbers. Although multiple global HMR methods have been proposed, they cannot faithfully capture climbing motions. To address the limitations of HMR methods for climbing, we propose Climbing-Cap, a motion recovery method that reconstructs continuous 3D human climbing motion in a global coordinate system. One key insight is to use the RGB and LiDAR modalities to separately reconstruct motions in camera coordinates and global coordinates and to optimize them jointly. We demonstrate the quality of the AscendMotion dataset and present promising results from ClimbingCap. The AscendMotion dataset and source code release publicly at http://www.lidarhumanmotion.net/climbingcap/
Xincheng Lin, Yuhua Luo, Shuqi Fan, Yudi Dai, Qixin Zhong, Lincai Zhong, Yuexin Ma, Lan Xu 0003, Chenglu Wen, Cheng Wang 0003
CVPR5
2024 RELI11D: A Comprehensive Multimodal Human Motion Dataset and Method
abstract
Comprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB, LiDAR, or IMU data. However, solely using these modalities or a combination of them may not be adequate for HPE, particularly for complex and fast movements. For holistic human motion understanding, we present RELI11D, a high-quality multimodal human motion dataset involves LiDAR, IMU system, RGB camera, and Event camera. It records the motions of 10 actors performing 5 sports in 7 scenes, including 3.32 hours of synchronized LiDAR point clouds, IMU measurement data, RGB videos and Event steams. Through extensive experiments, we demonstrate that the RELI 11 D presents considerable challenges and opportunities as it contains many rapid and complex motions that require precise location. To address the challenge of integrating different modalities, we propose LEIR, a multimodal baseline that effectively utilizes LiDAR Point Cloud, Event stream, and RGB through our cross-attention fusion strategy. We show that LEIR exhibits promising results for rapid motions and daily motions and that utilizing the characteristics of multiple modalities can indeed improve HPE performance. Both the dataset and source code release publicly in http://www.lidarhumanmotion.net/reli11d/, fostering collaboration and enabling further exploration in this field.
Shuqiang Cai, Shuqi Fan, Xincheng Lin, Yudi Dai, Chenglu Wen, Lan Xu 0003, Yuexin Ma, Cheng Wang 0003
CVPR6
2024 HmPEAR: A Dataset for Human Pose Estimation and Action Recognition
abstract
We introduce HmPEAR, a novel dataset crafted for advancing research in 3D Human Pose Estimation (3D HPE) and Human Action Recognition (HAR), with a primary focus on outdoor environments. This dataset offers a synchronized collection of imagery, LiDAR point clouds, 3D human poses, and action categories. In total, the dataset encompasses over 300,000 frames collected from 10 distinct scenes and 25 diverse subjects. Among these, 250,000 frames of data contain 3D human pose annotations captured using an advanced motion capture system and further optimized for accuracy. Furthermore, the dataset annotates 40 types of daily human actions, resulting in over 6,000 action clips. Through extensive experimentation, we have demonstrated the quality of HmPEAR and highlighted the challenges it presents to current methodologies. Additionally, we propose baselines leveraging sequential images and point clouds for 3D HPE and HAR, which underscore the mutual reinforcement between them, highlighting the potential for cross-task synergies. The dataset is available at http://www.lidarhumanmotion.net/hmpear.
Yitai Lin, Zhijie Wei, Wanfa Zhang, Xiping Lin, Yudi Dai, Chenglu Wen, Lan Xu 0003, Cheng Wang 0003
ACM Multimedia5
2024 HiSC4D: Human-Centered Interaction and 4D Scene Capture in Large-Scale Space Using Wearable IMUs and LiDAR
abstract
We introduce HiSC4D, a novel Human-centered interaction and 4D Scene Capture method, aimed at accurately and efficiently creating a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, rich human-human interactions, and human-environment interactions. By utilizing body-mounted IMUs and a head-mounted LiDAR, HiSC4D can capture egocentric human motions in unconstrained space without the need for external devices and pre-built maps. This affords great flexibility and accessibility for human-centered interaction and 4D scene capturing in various environments. Taking into account that IMUs can capture human spatially unrestricted poses but are prone to drifting for long-period using, and while LiDAR is stable for global localization but rough for local positions and orientations, HiSC4D employs a joint optimization method, harmonizing all sensors and utilizing environment cues, yielding promising results for long-term capture in large scenes. To promote research of egocentric human interaction in large scenes and facilitate downstream tasks, we also present a dataset, containing 8 sequences in 4 large scenes (200 to 5,000 [Formula: see text]), providing 36 k frames of accurate 4D human motions with SMPL annotations and dynamic scenes, 31k frames of cropped human point clouds, and scene mesh of the environment. A variety of scenarios, such as the basketball gym and commercial street, alongside challenging human motions, such as daily greeting, one-on-one basketball playing, and tour guiding, demonstrate the effectiveness and the generalization ability of HiSC4D. The dataset and code will be publicly available for research purposes.
Yudi Dai, Xiping Lin, Chenglu Wen, Lan Xu 0003, Yuexin Ma, Cheng Wang 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 SLOPER4D: A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments
abstract
We present SLOPER4D, a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild. Employing a head-mounted device integrated with a LiDAR and camera, we record 12 human subjects' activities over 10 diverse urban scenes from an egocentric view. Frame-wise annotations for 2D key points, 3D pose parameters, and global translations are provided, together with reconstructed scene point clouds. To obtain accurate 3D ground truth in such large dynamic scenes, we propose a joint optimization method to fit local SMPL meshes to the scene and fine-tune the camera calibration during dynamic motions frame by frame, resulting in plausible and scene-natural 3D human poses. Even-tually, SLOPER4D consists of 15 sequences of human motions, each of which has a trajectory length of more than 200 meters (up to 1,300 meters) and covers an area of more than 200 m2(up to 30,000 m2), including more than 100k LiDAR frames, 300k video frames, and 500k IMU-based motion frames. With SLOPER4D, we provide a detailed and thorough analysis of two critical tasks, including camera-based 3D HPE and LiDAR-based 3D HPE in urban environments, and benchmark a new task, GHPE. The in-depth analysis demonstrates SLOPER4D poses significant challenges to existing methods and produces great research opportunities. The dataset and code are released at http://www.lidarhumanmotion.net/sloper4d/.
Yudi Dai, Yitai Lin, Xiping Lin, Chenglu Wen, Lan Xu 0003, Hongwei Yi, Yuexin Ma, Cheng Wang 0003
CVPR1
2023 CIMI4D: A Large Multimodal Climbing Motion Dataset under Human-scene Interactions
abstract
Motion capture is a long-standing research problem. Although it has been studied for decades, the majority of research focus on ground-based movements such as walking, sitting, dancing, etc. Off- grounded actions such as climbing are largely overlooked. As an important type of action in sports and firefighting field, the climbing movements is challenging to capture because of its complex back poses, intricate human-scene interactions, and difficult global localization. The research community does not have an indepth understanding of the climbing action due to the lack of specific datasets. To address this limitation, we collect CIMI4D, a large rock Climbing Motion dataset from 12 persons climbing 13 different climbing walls. The dataset consists of around 180,000 frames of pose inertial measurements, LiDAR point clouds, RGB videos, high-precision static point cloud scenes, and reconstructed scene meshes. Moreover, we frame-wise annotate touch rock holds to facilitate a detailed exploration of human-scene interaction. The core of this dataset is a blending optimization process, which corrects for the pose as it drifts and is affected by the magnetic conditions. To evaluate the merit of CIMI4D, we perform four tasks which include human pose estimations (with/without scene constraints), pose prediction, and pose generation. The experimental results demonstrate that CIMI4D presents great challenges to existing methods and enables extensive research opportunities. We share the dataset with the research community in http://www.lidarhumanmotion.net/cimi4d/.
Yudi Dai, Chenglu Wen, Lan Xu 0003, Yuexin Ma, Cheng Wang 0003
CVPR3
2022 HSC4D: Human-centered 4D Scene Capture in Large-scale Indoor-outdoor Space Using Wearable IMUs and LiDAR
abstract
We propose Human-centered 4D Scene Capture (HSC4D) to accurately and efficiently create a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, and rich interactions between humans and environments. Using only body-mounted IMUs and LiDAR, HSC4D is space-free without any external devices' constraints and map-free without pre-built maps. Considering that IMUs can capture human poses but always drift for long-period use, while LiDAR is stable for global localization but rough for local positions and orientations, HSC4D makes both sensors complement each other by a joint optimization and achieves promising results for long-term capture. Relationships between humans and environments are also explored to make their interaction more realistic. To facilitate many down-stream tasks, like AR, VR, robots, autonomous driving, etc., we propose a dataset containing three large scenes (1k-5k m2) with accurate dynamic human motions and locations. Diverse scenarios (climbing gym, multi-story building, slope, etc.) and challenging human activities (exercising, walking up/down stairs, climbing, etc.) demonstrate the effectiveness and the generalization ability of HSC4D. The dataset and code is available at lidarhumanmotion.net/hsc4d.
Yudi Dai, Yitai Lin, Chenglu Wen, Lan Xu 0003, Jingyi Yu 0001, Yuexin Ma, Cheng Wang 0003
CVPR1
2022 Indoor 3D Human Trajectory Reconstruction Using Surveillance Camera Videos and Point Clouds
abstract
3D human trajectory reconstruction in an indoor scene is critical in various applications, such as indoor navigation and human activity recognition. This task is challenging due to occlusion and clutters of indoor scenes, flexible human body joints, and severe lack of relevant datasets. Although several methods have been proposed to reconstruct a 3D human trajectory, they either can recover only 2D positions or require human initiative cooperation. In this paper, we propose a novel framework for 3D human trajectory reconstruction in an indoor scene using monocular surveillance videos and static point clouds without any initiative cooperation. The proposed framework consists of three modules: 3D pose estimation, depth regression, and trajectory reconstruction. We first estimate 3D pose from videos. Especially, we reconstruct a half-body 3D pose to deal with the occlusion problem. Then, we propose a depth regression approach to iteratively regress the depth of a 3D pose. Unlike data-driven approaches, our depth regression approach does not require training data and can be integrated into any 3D pose model. Finally, we exploit the geometric constraints from the point cloud to optimize the 3D trajectory. We evaluated the 3D pose estimation and depth regression modules on the H3.6M datasets. Due to the lack of evaluation datasets, we also built a trajectory dataset to evaluate the trajectory reconstruction performance. Empirical evaluation shows that our framework achieves accurate trajectory reconstruction results on real-world videos.
Yudi Dai, Chenglu Wen, Yulan Guo, Longbiao Chen, Cheng Wang 0003
IEEE Trans. Circuits Syst. Video Technol.1
2020 Toward Efficient 3-D Colored Mapping in GPS-/GNSS-Denied Environments
abstract
Efficient 3-D mapping provides useful and detailed 3-D data for many applications. In this letter, we present a multisensor calibration and mapping method, to provide highly efficient and relatively accurate colored mapping for GPS-/global navigation satellite system-denied environments. The sensor data include 3-D laser scanning point clouds and camera images. A simultaneous localization and mapping (SLAM)-assisted calibration method is first proposed for multiple multibeam light detection and ranging (LiDAR) and multiple camera calibration. An improved SLAM method with loop closure is proposed for 3-D mapping. With the proposed calibration and mapping methods, centimeter-level colored point clouds can be obtained efficiently. The proposed method was tested with both backpacked and car-mounted systems on indoor and outdoor scenes. Experimental results show the effectiveness and efficiency of the proposed calibration and mapping methods.
Chenglu Wen, Yudi Dai, Yan Xia 0003, Yuhan Lian, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2020 Urban 3D modeling with mobile laser scanning: a review
abstract
Mobile laser scanning (MLS) systems mainly comprise laser scanners and mobile mapping platforms. Typical MLS systems are able to acquire three-dimensional point clouds with 1-10 centimeter point spacing at a normal driving or walking speed in the street or indoor environments. The MLS' advantages of efficiency and stability make it a quite practical tool for three-dimensional urban modeling. This paper reviews the latest advances in 3D modeling of the LiDAR-based mobile mapping system (MMS) point cloud, including LiDAR Simultaneous Localization and Mapping (SLAM), point cloud registration, feature extraction, object extraction, semantic segmentation, and deep learning processing. Then typical urban modeling applications based on MMS are also discussed.
Cheng Wang 0003, Chenglu Wen, Yudi Dai, Shangshu Yu, Minghao Liu 0007
Virtual Real. Intell. Hardw.3
2018 Line Structure-Based Indoor and Outdoor Integration Using Backpacked and TLS Point Cloud Data
abstract
This letter presents a line structure-based method for integration of centimeter-level indoor backpacked scanning point clouds and millimeter-level outdoor terrestrial laser scanning point clouds. Using 3-D lines for registration, instead of matching points directly, can improve the robustness of the method and adapt to multisource point cloud data of different qualities. Considering the limited overlapping between indoor and outdoor scenes, line structures are extracted from overlapped wall areas that may be included in interior and exterior data. Here, a patch-based method labels a point cloud into wall, ceiling, floor categories, as well as assigning the candidate overlapping walls. Then, lines structures are extracted from the wall plane point cloud. Potential door and window line structures are detected and refined for point cloud registration. Last, an iterative closest point-based method is used to fine tune the registration results. Our results show that the proposed method effectively integrates a promising map of indoor and outdoor scenes.
Chenglu Wen, Xiaotian Sun 0005, Shiwei Hou, Jinbin Tan, Yudi Dai, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5