Yichen Yao 0001

dblp:178/7287-1 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0006-9319-9339ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OptimalCap: Efficient and Robust LiDAR-Based Motion Capture in Free Environments
abstract
LiDAR-based human motion capture holds great promise for large-scale, unconstrained environments. However, existing approaches often rely on clean, pre-segmented point clouds and struggle with noisy or dynamic scenes, limiting their practical applicability. We propose OptimalCap, a robust and efficient LiDAR-based framework that integrates hierarchical skeletal modeling and kinematic-aware temporal optimization to enable accurate, coherent, and real-time multi-human motion capture. To support training and evaluation under realistic disturbances, we also introduce NoiseMotion, a large-scale synthetic dataset simulating human-object interactions in noisy environments. Extensive experiments on public and synthetic benchmarks demonstrate that OptimalCap achieves state-of-the-art accuracy, robustness, and temporal consistency, while supporting over 20 individuals, at 60 FPS and up to 100 meters, setting a new standard for scalable, real-world LiDAR-based motion capture.
Yiming Ren 0001, Yujing Sun 0001, Yichen Yao 0001, Xiaoxiao Long, Xinge Zhu, Siu-Ming Yiu, Yuexin Ma
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Can LVLMs Obtain a Driver's License? A Benchmark Towards Reliable AGI for Autonomous Driving
abstract
Large Vision-Language Models (LVLMs) have recently garnered significant attention, with many efforts aimed at harnessing their general knowledge to enhance the interpretability and robustness of autonomous driving models. However, LVLMs typically rely on large, general-purpose datasets and lack the specialized expertise required for professional and safe driving. Existing vision-language driving datasets focus primarily on scene understanding and decision-making, without providing explicit guidance on traffic rules and driving skills, which are critical aspects directly related to driving safety. To bridge this gap, we propose IDKB, a large-scale dataset containing over one million data items collected from various countries, including driving handbooks, theory test data, and simulated road test data. Much like the process of obtaining a driver's license, IDKB encompasses nearly all the explicit knowledge needed for driving from theory to practice. In particular, we conducted comprehensive tests on 15 LVLMs using IDKB to assess their reliability in the context of autonomous driving and provided extensive analysis. We also fine-tuned popular models, achieving notable performance improvements, which further validate the significance of our dataset.
Yichen Yao 0001, Jiadong Tu, Jiangnan Shao, Yuexin Ma, Xinge Zhu
AAAI2
2025 STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
abstract
The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature misalignment due to inadequate decoupling of spatio-temporal dynamics and limited cross-frame feature propagation mechanisms. To address these limitations, we present STAGE (Streaming Temporal Attention Generative Engine), a novel auto-regressive framework that pioneers hierarchical feature coordination and multi-phase optimization for sustainable video synthesis. To achieve high-quality long-horizon driving video generation, we introduce Hierarchical Temporal Feature Transfer (HTFT) and a novel multi-stage training strategy. HTFT enhances temporal consistency between video frames throughout the video generation process by modeling the temporal and denoising process separately and transferring denoising features between frames. The multi-stage training strategy is to divide the training into three stages, through model decoupling and auto-regressive inference process simulation, thereby accelerating model convergence and reducing error accumulation. Experiments on the Nuscenes dataset show that STAGE has significantly surpassed existing methods in the long-horizon driving video generation task. In addition, we also explored STAGE’s ability to generate unlimited-length driving videos. We generated 600 frames of high-quality driving videos on the Nuscenes dataset, which far exceeds the maximum length achievable by existing methods. Our project homepage:https://4dvlab.github.io/STAGE/
Yichen Yao 0001, Yaming Wang, Qingqiu Huang, Yuexin Ma, Xinge Zhu
IROS2
2024 HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real Scenes
abstract
Human-centric 3D scene understanding has recently drawn increasing attention, driven by its critical impact on robotics. However, human-centric real-life scenarios are extremely diverse and complicated, and humans have intri-cate motions and interactions. With limited labeled data, supervised methods are difficult to generalize to general scenarios, hindering real-life applications. Mimicking human intelligence, we propose an unsupervised 3D detection method for human-centric scenarios by transferring the knowledge from synthetic human instances to real scenes. To bridge the gap between the distinct data representations and feature distributions of synthetic models and real point clouds, we introduce novel modules for effective instance-to-scene representation transfer and synthetic-to-real feature alignment. Remarkably, our method exhibits superior performance compared to current state-of-the-art techniques, achieving 87.8% improvement in mAP and closely approaching the performance of fully supervised methods (62.15 mAP vs. 69.02 mAP) on HuCenLife Dataset.
Yichen Yao 0001, Zimo Jiang, Yujing Sun 0001, Zhencai Zhu, Xinge Zhu, Runnan Chen, Yuexin Ma
CVPR1
2024 LiveHPS++: Robust and Coherent Motion Capture in Dynamic Free Environment
Yiming Ren 0001, Yichen Yao 0001, Xiaoxiao Long, Yujing Sun 0001, Yuexin Ma
ECCV (29)3
2024 RealDex: Towards Human-like Grasping for Robotic Dexterous Hand
Yaxun Yang, Youzhuo Wang, Yichen Yao 0001, Sören Schwertfeger, Sibei Yang, Wenping Wang 0001, Jingyi Yu 0001, Xuming He 0001, Yuexin Ma
IJCAI6
2024 Towards Practical Human Motion Prediction with LiDAR Point Clouds
abstract
Human motion prediction is crucial for human-centric multimedia understanding and interacting. Current methods typically rely on ground truth human poses as observed input, which is not practical for real-world scenarios where only raw visual sensor data is available. To implement these methods in practice, a pre-phrase of pose estimation is essential. However, such two-stage approaches often lead to performance degradation due to the accumulation of errors. Moreover, reducing raw visual data to sparse keypoint representations significantly diminishes the density of information, resulting in the loss of fine-grained features. In this paper, we propose LiDAR-HMP, the first single-LiDAR-based 3D human motion prediction approach, which receives the raw LiDAR point cloud as input and forecasts future 3D human poses directly. Building upon our novel structure-aware body feature descriptor, LiDAR-HMP adaptively maps the observed motion manifold to future poses and effectively models the spatial-temporal correlations of human motions for further refinement of prediction results. Extensive experiments show that our method achieves state-of-the-art performance on two public benchmarks and demonstrates remarkable robustness and efficacy in real-world deployments. https://4dvlab.github.io/project_page/LiDARHMP.html
Yiming Ren 0001, Yichen Yao 0001, Yujing Sun 0001, Yuexin Ma
ACM Multimedia3
2023 Human-centric Scene Understanding for 3D Large-scale Scenarios
abstract
Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset for human-centric scene under-standing, dubbed HuCenLife, which is collected in diverse daily-life scenarios with rich and fine-grained annotations. Our HuCenLife can benefit many 3D perception tasks, such as segmentation, detection, action recognition, etc., and we also provide benchmarks for these tasks to facilitate related research. In addition, we design novel modules for LiDAR-based segmentation and action recognition, which are more applicable for large-scale human-centric scenarios and achieve state-of-the-art performance. The dataset and code can be found at https://github.com/4DVLab/HuCenLife.git.
Yiteng Xu, Peishan Cong, Yichen Yao 0001, Runnan Chen, Yuenan Hou, Xinge Zhu, Xuming He 0001, Jingyi Yu 0001, Yuexin Ma
ICCV3