Yining Shi 0002

dblp:161/3927-2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-2926-925XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion Prediction
abstract
LiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion
Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang
IEEE Trans. Intell. Transp. Syst.5
2025 PriorMotion: Generative Class-Agnostic Motion Prediction with Raster-Vector Motion Field Priors
abstract
Reliable spatial and motion perception is essential for safe autonomous navigation. Recently, class-agnostic motion prediction on bird's-eye view (BEV) cell grids derived from LiDAR point clouds has gained significant attention. However, existing frameworks typically perform cell classification and motion prediction on a per-pixel basis, neglecting important motion field priors such as rigidity constraints, temporal consistency, and future interactions between agents. These limitations lead to degraded performance, particularly in sparse and distant regions. To address these challenges, we introduce \textbf{PriorMotion}, an innovative generative framework designed for class-agnostic motion prediction that integrates essential motion priors by modeling them as distributions within a structured latent space. Specifically, our method captures structured motion priors using raster-vector representations and employs a variational autoencoder with distinct dynamic and static components to learn future motion distributions in the latent space. Experiments on the nuScenes dataset demonstrate that \textbf{PriorMotion} outperforms state-of-the-art methods across both traditional metrics and our newly proposed evaluation criteria. Notably, we achieve improvements of approximately 15.24\% in accuracy for fast-moving objects, an 3.59\% increase in generalization, a reduction of 0.0163 in motion stability, and a 31.52\% reduction in prediction errors in distant regions. Further validation on FMCW LiDAR sensors confirms the robustness of our approach.
Kangan Qian, Jinyu Miao, Xinyu Jiao, Ziang Luo, Zheng Fu, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Diange Yang
ICCV6
2025 SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
abstract
The well-established modular autonomous driving system is decoupled into different standalone tasks, e.g. perception, prediction and planning, suffering from information loss and error accumulation across modules. In contrast, end-to-end paradigms unify multi-tasks into a fully differentiable framework, allowing for optimization in a planning-oriented spirit. Despite the great potential of end-to-end paradigms, both the performance and efficiency of existing methods are not satisfactory, particularly in terms of planning safety. We attribute this to the computationally expensive BEV (bird's eye view) features and the straightforward design for prediction and planning. To this end, we explore the sparse representation and review the task design for end-to-end autonomous driving, proposing a new paradigm named SparseDrive. Concretely, SparseDrive consists of a symmetric sparse perception module and a parallel motion planner. The sparse perception module unifies detection, tracking and online mapping with a symmetric model architecture, learning a fully sparse representation of the driving scene. For motion prediction and planning, we review the great similarity between these two tasks, leading to a parallel design for motion planner. Based on this parallel design, which models planning as a multi-modal problem, we propose a hierarchical planning selection strategy, which incorporates a collision-aware rescore module, to select a rational and safe trajectory as the final planning output. With such effective designs, SparseDrive surpasses previous state-of-the-arts by a large margin in performance of all tasks, while achieving much higher training and inference efficiency.
Xuewu Lin, Yining Shi 0002, Sifa Zheng
ICRA3
2025 LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction
abstract
Accurate spatial and motion understanding is critical for autonomous driving systems. While object-level perception models excel in structured environments, they struggle with open-set categories and often lack precise geometric representation. Occupancy-based, class-agnostic methods offer better scene expressiveness but typically ignore inter-agent interactions and fail to ensure physical consistency in motion predictions, limiting their reliability in complex traffic scenarios. In this paper, we propose LEGO-Motion, a novel class-agnostic motion prediction framework that bridges the gap between instance-level reasoning and occupancy-based modeling. Unlike conventional grid-based methods that treat each cell independently, LEGO-Motion introduces two key components: (1) the Interaction-Augmented Instance Encoder (IaIE), which models interactions among dynamic agents via cross-attention, and (2) the Instance-Enhanced BEV Encoder (IeBE), which improves motion consistency across instances through multi-stage feature fusion. These components enable our model to learn semantically coherent and physically plausible motion fields. Extensive experiments on the nuScenes dataset show that LEGO-Motion achieves a around 6% improvement in motion prediction accuracy over the previous state-of-the-art, while maintaining real-time inference at 21ms. Moreover, our method demonstrates strong generalization on a proprietary FMCW LiDAR benchmark. These results validate LEGO-Motion's effectiveness in capturing both global scene structure and fine-grained motion dynamics, making it a promising foundation for next-generation perception systems.
Kangan Qian, Jinyu Miao, Ziang Luo, Zheng Fu, Jinchen Li, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang
IROS6
2025 EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Driving
abstract
3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or transformers at the voxel level. Moreover, OccNets are supervised with expensive large-scale dense voxel labels. Model and label inefficiencies, caused by excessive network parameters and label annotation requirements, severely hinder the onboard deployment of OccNets. This paper proposes an EFFicient Occupancy learning framework, EFFOcc, that targets minimal network complexity and label requirements while achieving state-of-the-art accuracy. We first propose an efficient fusion-based OccNet that only uses simple 2D operators and improves accuracy to the state-of-the-art on three large-scale benchmarks: Occ3D-nuScenes, Occ3D-Waymo, and OpenOccupancy-nuScenes. On the Occ3D-nuScenes benchmark, the fusion-based model with ResNet-18 as the image backbone has 21.35M parameters and achieves 51.49 in terms of mean Intersection over Union (mIoU). Furthermore, we propose a multi-stage occupancy-oriented distillation to efficiently transfer knowledge to vision-only OccNet. Extensive experiments on occupancy benchmarks show state-of-the-art precision for both fusion-based and vision-based OccNets. For the demonstration of learning with limited labels, we achieve 94.38% of the performance (mIoU = 28.38) of a 100% labeled vision OccNet (mIoU = 30.07) using the same OccNet trained with only 40% labeled sequences and distillation from the fusion-based OccNet. Code is available at https://github.com/synsin0/EFFOcc.
Yining Shi 0002, Kun Jiang 0002, Jinyu Miao, Ke Wang 0021, Kangan Qian, Yunlong Wang 0009, Jiusi Li, Tuopu Wen, Mengmeng Yang 0001, Yiliang Xu, Diange Yang
IROS1
2025 COME: Adding Scene-Centric Forecasting Control to Occupancy World Model
abstract
World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to separate environmental changes from ego-motion by leveraging the scene-centric coordinate systems. In this paper, we introduce COME: a framework that integrates scene-centric forecasting Control into the Occupancy world ModEl. Specifically, COME first generates ego-irrelevant, spatially consistent future features through a scene-centric prediction branch, which are then converted into scene condition using a tailored ControlNet. These condition features are subsequently injected into the occupancy world model, enabling more accurate and controllable future occupancy predictions. Experimental results on the nuScenes-Occ3D dataset show that COME achieves consistent and significant improvements over state-of-the-art (SOTA) methods across diverse configurations, including different input sources (ground-truth, camera-based, fusion-based occupancy) and prediction horizons (3s and 8s). For example, under the same settings, COME achieves 26.3% better mIoU metric than DOME and 23.7% better mIoU metric than UniScene. These results highlight the efficacy of disentangled representation learning in enhancing spatio-temporal prediction fidelity for world models. Code is available at https://github.com/synsin0/COME.
Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Tuopu Wen, Mengmeng Yang 0001, Diange Yang
NeurIPS1
2025 Grid-Centric Traffic Scenario Perception for Autonomous Driving: A Comprehensive Review
abstract
The grid-centric perception is a crucial field for mobile robot perception and navigation. Nonetheless, the grid-centric perception is less prevalent than object-centric perception as autonomous vehicles need to accurately perceive highly dynamic, large-scale traffic scenarios, and the complexity and computational costs of grid-centric perception are high. In recent years, the rapid development of deep learning techniques and hardware provides fresh insights into the evolution of grid-centric perception. The fundamental difference between grid-centric and object-centric pipeline lies in that grid-centric perception follows a geometry-first paradigm which is more robust to the open-world driving scenarios with endless long-tailed semantically unknown obstacles. Recent research demonstrates the great advantages of grid-centric perception, such as comprehensive fine-grained environmental representation, greater robustness to occlusion and irregular-shaped objects, better ground estimation, and safer planning policies. There is also a growing trend that the capacity of occupancy networks is greatly expanded to 4-D scene perception and prediction, and the latest techniques are highly related to new research topics, such as 4-D occupancy forecasting, generative artificial intelligence (GenAI), and world models in the field of autonomous driving. Given the lack of current surveys for this rapidly expanding field, we present a hierarchically structured review of grid-centric perception for autonomous vehicles. We organize previous and current knowledge of occupancy grid techniques along the main vein from 2-D bird-eye view (BEV) grids to 3-D occupancy to 4-D occupancy forecasting. We additionally summarize label-efficient occupancy learning and the role of grid-centric perception in driving systems. Finally, we present a summary of the current research trend and provide future outlooks.
Yining Shi 0002, Kun Jiang 0002, Jiusi Li, Zelin Qian, Junze Wen, Mengmeng Yang 0001, Ke Wang 0021, Diange Yang
IEEE Trans. Neural Networks Learn. Syst.1
2024 PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous Driving
abstract
Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect obstacles regardless of their shape and occlusion. Modern occupancy networks mainly focus on reconstructing visible voxels from object surfaces with voxel-wise semantic prediction. Usually, they suffer from inconsistent predictions of one object and mixed predictions for adjacent objects. These confusions may harm the safety of downstream planning modules. To this end, we investigate panoptic segmentation on 3D voxel scenarios and propose an instance-aware occupancy network, PanoSSC. We predict foreground objects and backgrounds separately and merge both in post-processing. For foreground instance grouping, we propose a novel 3D instance mask decoder that can efficiently extract individual objects. we unify geometric reconstruction, 3D semantic segmentation, and 3D instance segmentation into PanoSSC framework and propose new metrics for evaluating panoptic voxels. Extensive experiments show that our method achieves competitive results on SemanticKITTI semantic scene completion benchmark.
Yining Shi 0002, Jiusi Li, Kun Jiang 0002, Ke Wang 0021, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang
3DV1
2024 StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation
abstract
Predicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However, current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict uniform snapshots of future occupancy states and require strictly synchronized sensory data for sensor fusion. We propose a novel framework, StreamingFlow, to lift these strong limitations. StreamingFlow is a novel BEV occupancy predictor that ingests asynchronous multi-sensor data streams for fusion and performs streaming fore-casting of the future occupancy map at any future times-tamps. By integrating neural ordinary differential equations (N-ODE) into recurrent neural networks, StreamingFlow learns derivatives of BEV features over temporal horizons, updates the implicit sensor's BEV features as part of the fusion process, and propagates BEV states to the desired future time point. It shows good zero-shot generalization ability of prediction, reflected in the interpolation of the ob-served prediction time horizon and the reasonable inference of the unseen farther future period. Extensive experiments on two large-scale datasets, nuScenes [2] and Lyft L5 [14], demonstrate that StreamingFlow significantly outperforms previous vision-based, LiDAR-based methods, and shows superior performance compared to state-of-the-art fusion-based methods.
Yining Shi 0002, Kun Jiang 0002, Ke Wang 0021, Jiusi Li, Yunlong Wang 0009, Mengmeng Yang 0001, Diange Yang
CVPR1