Peidong Li

dblp:206/2707 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 38% Autonomous driving · 38% Video understanding and tracking · 24%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › Video understanding and tracking › object tracking
3d object tracking
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Robotics › Autonomous driving › perception
3d perception
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › Video understanding and tracking
multi-object tracking
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision
spatio-temporal alignment
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision › 3d object detection
temporal 3d detection
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Robotics › Autonomous driving
end-to-end driving
0.912025
Navigation-Guided Sparse Scene Representation for End-to-End Autonomous Driving · ICLR 2025
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception
0.812024
DualBEV: Unifying Dual View Transformation with Probabilistic Correspondences · ECCV (84) 2024
Robotics › Autonomous driving
perception
0.312025
Navigation-Guided Sparse Scene Representation for End-to-End Autonomous Driving · ICLR 2025
Computer vision › 3D vision
view transformation
0.212024
DualBEV: Unifying Dual View Transformation with Probabilistic Correspondences · ECCV (84) 2024

Methods — techniques the papers use, named apart from their topics

multi-hypothesis decoding · 1.0motion model · 1.0attention mechanism · 1.0temporal enhancement · 0.9self-supervised learning · 0.9probabilistic correspondence · 0.8
YearPublicationVenuePosition
2026 Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
abstract
Spatio-temporal alignment is crucial for temporal modeling of end-to-end (E2E) perception in autonomous driving (AD), providing valuable structural and textural prior information. Existing methods typically rely on the attention mechanism to align objects across frames, simplifying the motion model with a unified explicit physical model (constant velocity, etc.). These approaches prefer semantic features for implicit alignment, challenging the importance of explicit motion modeling in the traditional perception paradigm. However, variations in motion states and object features across categories and frames render this alignment suboptimal. To address this, we propose HAT, a spatio-temporal alignment module that allows each object to adaptively decode the optimal alignment proposal from multiple hypotheses without direct supervision. Specifically, HAT first utilizes multiple explicit motion models to generate spatial anchors and motion-aware feature proposals for historical instances. It then performs multi-hypothesis decoding by incorporating semantic and motion cues embedded in cached object queries, ultimately providing the optimal alignment proposal for the target frame. On nuScenes, HAT consistently improves 3D temporal detectors and trackers across diverse baselines. It achieves state-of-the-art tracking results with 46.0% AMOTA on the test set when paired with the DETR3D detector. In an object-centric E2E AD method, HAT enhances perception accuracy (+1.3% mAP, +3.1% AMOTA) and reduces the collision rate by 32%. When semantics are corrupted (nuScenes-C), the enhancement of motion modeling by HAT enables more robust perception and planning in the E2E AD.
Peidong Li, Dedong Liu, Jiajia Fu, Dixiao Cui, Lijun Zhao 0003, Lining Sun
AAAI2
2025 Navigation-Guided Sparse Scene Representation for End-to-End Autonomous Driving
abstract
End-to-End Autonomous Driving (E2EAD) methods typically rely on supervised perception tasks to extract explicit scene information (e.g., objects, maps). This reliance necessitates expensive annotations and constrains deployment and data scalability in real-time applications. In this paper, we introduce SSR, a novel framework that utilizes only 16 navigation-guided tokens as Sparse Scene Representation, efficiently extracting crucial scene information for E2EAD. Our method eliminates the need for human-designed supervised sub-tasks, allowing computational resources to concentrate on essential elements directly related to navigation intent. We further introduce a temporal enhancement module, aligning predicted future scenes with actual future scenes through self-supervision. SSR achieves a 27.2\% relative reduction in L2 error and a 51.6\% decrease in collision rate to UniAD in nuScenes, with a 10.9× faster inference speed and 13× faster training time. Moreover, SSR outperforms VAD-Base with a 48.6-point improvement on driving score in CARLA's Town05 Long benchmark. This framework represents a significant leap in real-time autonomous driving systems and paves the way for future scalable deployment. Code is available at https://github.com/PeidongLi/SSR.
Peidong Li, Dixiao Cui
ICLR1
2025 Power-GNN: a graph over-sampling method to mitigate power-law distribution in graph neural networks
Peidong Li, Zhenghong Zhong, Yangguang Zhao, Changheng Shao, Yi Sui 0003, Rencheng Sun
Appl. Intell.1
2024 DualBEV: Unifying Dual View Transformation with Probabilistic Correspondences
Peidong Li, Wancheng Shen, Qihao Huang, Dixiao Cui
ECCV (84)1