Shouyi Lu

dblp:293/0812 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-5055-1802ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 47% Robot navigation and mapping · 32% Vision and language · 11%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
point cloud processing
1.922026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025
Computer vision › 3D vision
3d object detection
1.012026
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision › point cloud processing › radar point cloud processing
4d radar point cloud
1.012026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
Robotics › Robot navigation and mapping › localization
odometry
1.012026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
Robotics › Robot navigation and mapping › localization › odometry
radar odometry
1.012026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
Computer vision › Vision and language
vision-language dataset
1.012026
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Video understanding and tracking
spatiotemporal aggregation
0.912025
TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025
Robotics › Robot navigation and mapping › place recognition
visual place recognition
0.912025
TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.312026
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026

Methods — techniques the papers use, named apart from their topics

surround-view camera · 1.0gauss-newton optimization · 1.0dual-stream backbone · 1.0differentiable neural-optimization iteration · 1.0LiDAR · 1.04d imaging radar · 1.0trajectory-guided alignment · 0.9deformable feature aggregation · 0.9
YearPublicationVenuePosition
2026 DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations
abstract
A novel learning-optimization-combined 4D radar odometry model, named DNOI-4DRO, is proposed in this paper. The proposed model seamlessly integrates traditional geometric optimization with end-to-end neural network training, leveraging an innovative differentiable neural-optimization iteration operator. In this framework, point-wise motion flow is first estimated using a neural network, followed by the construction of a cost function based on the relationship between point motion and pose in 3D space. The radar pose is then refined using Gauss-Newton updates. Additionally, we design a dual-stream 4D radar backbone that integrates multi-scale geometric features and clustering-based class-aware features to enhance the representation of sparse 4D radar point clouds. Extensive experiments on the VoD and Snail-Radar datasets demonstrate the superior performance of our model, which outperforms recent classical and learning-based approaches. Notably, our method even achieves results comparable to A-LOAM with mapping optimization using LiDAR point clouds as input.
Shouyi Lu, Huanyu Zhou, Guirong Zhuo
AAAI1
2026 Diff-GNSS: Diffusion-Based GNSS Pseudorange Error Estimation for Accurate Positioning
Shouyi Lu, Ziyao Li, Guirong Zhuo, Lu Xiong 0001
IEEE Internet Things J.2
2026 OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
abstract
The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions. Next-generation autonomous driving datasets must be multimodal, incorporating data from advanced sensors that feature extensive data coverage, detailed annotations, and diverse scene representation. To address this need, we present OmniHD-Scenes, a large-scale multimodal dataset that provides comprehensive omnidirectional high-definition data. The OmniHD-Scenes dataset combines data from 128-beam LiDAR, six cameras, and six 4D imaging radar systems to achieve full environmental perception. The dataset comprises 1501 clips, each approximately 30-s long, totaling more than 450 K synchronized frames and more than 5.85 million synchronized sensor data points. We also propose a novel 4D annotation pipeline. To date, we have annotated 200 clips with more than 514 K precise 3D bounding boxes. These clips also include semantic segmentation annotations for static scene elements. Additionally, we introduce a novel automated pipeline for generation of the dense occupancy ground truth, which effectively leverages information from non-key frames. Alongside the proposed dataset, we establish comprehensive evaluation metrics, baseline models, and benchmarks for 3D detection and semantic occupancy prediction. These benchmarks utilize surround-view cameras and 4D imaging radar to explore cost-effective sensor solutions for autonomous driving applications. Extensive experiments demonstrate the effectiveness of our low-cost sensor configuration and its robustness under adverse conditions.
Lianqing Zheng, Qunshu Lin, Wenjin Ai, Minghao Liu 0021, Shouyi Lu, Hongze Ren, Jingyue Mo, Xiaokai Bai, Zhixiong Ma, Xichan Zhu
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Doracamom: Joint 3D Detection and Occupancy Prediction With Multi-View 4D Radars and Cameras for Omnidirectional Perception
Lianqing Zheng, Runwei Guan, Shouyi Lu, Xiaokai Bai, Zhixiong Ma, Xichan Zhu
IEEE Trans. Circuits Syst. Video Technol.5
2025 TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition
abstract
Place recognition is essential for achieving closedloop or global positioning in autonomous vehicles and mobile robots. Despite recent advancements in place recognition using 2D cameras or 3D LiDAR, it remains to be seen how to use 4D radar for place recognition - an increasingly popular sensor for its robustness against adverse weather and lighting conditions. Compared to LiDAR point clouds, radar data are drastically sparser, noisier and in much lower resolution, which hampers their ability to effectively represent scenes, posing significant challenges for 4D radar-based place recognition. This work addresses these challenges by leveraging multimodal information from sequential 4D radar scans and effectively extracting and aggregating spatio-temporal features. Our approach follows a principled pipeline that comprises (1) dynamic points removal and ego-velocity estimation from velocity property, (2) bird's eye view (BEV) feature encoding on the refined point cloud, (3) feature alignment using BEV feature map motion trajectory calculated by ego-velocity, (4) multiscale spatio-temporal features of the aligned BEV feature maps are extracted and aggregated. Real-world experimental results validate the feasibility of the proposed method and demonstrate its robustness in handling dynamic environments. Source codes are available.
Shouyi Lu, Guirong Zhuo, Huanyu Zhou, Renbo Huang, Minqing Huang, Lianqing Zheng, Qiang Shu
ICRA1
2025 R2LDM: An Efficient 4D Radar Super-Resolution Framework Leveraging Diffusion Model
abstract
We introduce R2LDM, an innovative approach for generating dense and accurate 4D radar point clouds, guided by corresponding LiDAR point clouds. Instead of utilizing range images or bird’s eye view (BEV) images, we represent both LiDAR and 4D radar point clouds using voxel features, which more effectively capture 3D shape information. Subsequently, we propose the Latent Voxel Diffusion Model (LVDM), which performs the diffusion process in the latent space. Additionally, a novel Latent Point Cloud Reconstruction (LPCR) module is utilized to reconstruct point clouds from high-dimensional latent voxel features. As a result, R2LDM effectively generates LiDAR-like point clouds from paired raw radar data. We evaluate our approach on two different datasets, and the experimental results demonstrate that our model achieves 6- to 10-fold densification of radar point clouds, outperforming state-of-the-art baselines in 4D radar point cloud super-resolution. Furthermore, the enhanced radar point clouds generated by our method significantly improve downstream tasks, achieving up to 31.7% improvement in point cloud registration recall rate and 24.9% improvement in object detection accuracy.
Shouyi Lu, Renbo Huang, Minqing Huang, Wei Tian 0001, Guirong Zhuo, Lu Xiong 0001
IROS2
2025 4DRE-VIO: 4D Radar-Enhanced Monocular Visual Inertial Odometry for Urban Environments
abstract
Localization is crucial for Intelligent Transportation Systems (ITS) as it provides the precise position and orientation necessary for the operation of autonomous vehicles. Monocular vision-inertial odometry (VIO) has gained extensive application due to its high accuracy and cost-effectiveness. In this paper, we propose 4DRE-VIO, a novel 4D radar-enhanced monocular VIO system that achieves accurate and cost-effective pose estimation for autonomous vehicles in urban environments. By tightly integrating 4D radar measurements, we address the issues of scale ambiguity, scale drift, and dynamic object interference associated with monocular VIO. Specifically, we integrate 4D radar Doppler velocity with non-holonomic constraints (NHC) to provide a reliable velocity observation. Moreover, this velocity is directly used to construct a monocular scale observation, ensuring stable and accurate scale estimation. To mitigate the impact of dynamic objects on VIO, which is based on static environment assumptions, we introduce a 4D radar-enhanced visual semantic segmentation algorithm for dynamic object recognition. Additionally, we utilize a factor graph to tightly integrate measurements from the 4D radar, IMU, and monocular camera. The design of adjacent factors and optimal co-visible factors ensures the efficient use of visual information. An adaptive fusion strategy based on driving conditions maximizes the benefits of each sensor. We validate the effectiveness of the proposed method through extensive testing on a large dataset collected in urban environments, including open streets and residential districts. Compared to typical VIO and 4D radar-related odometry, our method demonstrates superior performance on our dataset and the public dataset.
Wufei Fu, Shouyi Lu, Guirong Zhuo, Lu Xiong 0001
IEEE Trans. Intell. Transp. Syst.3