Xiangyuan Peng

dblp:310/7579 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0007-0742-8404ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 50% Representation and self-supervised learning · 25% Probabilistic and Bayesian machine learning · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent world model
0.912025
Vector Quantization in the Brain: Grid-like Codes in World Models · NeurIPS 2025
Machine learning › Reinforcement learning
model-based reinforcement learning
0.912025
Vector Quantization in the Brain: Grid-like Codes in World Models · NeurIPS 2025
Machine learning › Representation and self-supervised learning
vector quantization
0.912025
Vector Quantization in the Brain: Grid-like Codes in World Models · NeurIPS 2025
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.912025
Vector Quantization in the Brain: Grid-like Codes in World Models · NeurIPS 2025
Emerging computing paradigms › neuromorphic computing
attractor neural network
0.312025
Vector Quantization in the Brain: Grid-like Codes in World Models · NeurIPS 2025
Emerging computing paradigms
neuromorphic computing
0.312025
Vector Quantization in the Brain: Grid-like Codes in World Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

vector quantization · 1.7continuous attractor neural network · 1.7
YearPublicationVenuePosition
2026 Text4Radar-V2X: Text-guided 4D Radar for Cooperative 3D Object Detection
abstract
Vehicle-to-Everything (V2X) perception enhances 3D object detection by extending sensing range and mitigating occlusions through information sharing between infrastructure- and vehicle-mounted sensors. Among various sensing modalities, 4D radar has attracted increasing attention for V2X perception due to its ability to provide 3D point clouds and velocity measurements, as well as its robustness under adverse weather. However, 4D radar point clouds remain sparse and noisy. Recent advances in vision-language models (VLMs) have enabled high-level scene understanding from visual inputs, with strong generalization to complex and unseen scenes. Motivated by this, we propose a novel 4D radar and text fusion framework, Text4Radar-V2X, which leverages text semantics to compensate for the sparsity of 4D radar features. Specifically, we introduce a view-specific asymmetric text-generation strategy. The generated Q&A pairs contain background structural semantics from infrastructure perspectives and foreground object semantics from vehicle perspectives. Furthermore, we design a dual-branch text-driven interaction to hierarchically integrate asymmetric text with 4D radar point clouds. Extensive experiments on the V2X-R dataset demonstrate that our method achieves the best mAP at IoU thresholds of 0.3 and 0.5, with improvements of 2.52% and 2.07%, respectively.
Xiangyuan Peng, Kay Bierzynski, Lorenzo Servadei, Robert Wille
ICMR1
2025 MutualForce: Mutual-Aware Enhancement for 4D Radar-LiDAR 3D Object Detection
abstract
Radar and LiDAR have been widely used in autonomous driving as LiDAR provides rich structure information, and radar demonstrates high robustness under adverse weather. Recent studies highlight the effectiveness of fusing radar and LiDAR point clouds. However, challenges remain due to the modality misalignment and information loss during feature extractions. To address these issues, we propose a 4D radar-LiDAR framework to mutually enhance their representations. Initially, the indicative features from radar are utilized to guide both radar and LiDAR geometric feature learning. Subsequently, to mitigate their sparsity gap, the shape information from LiDAR is used to enrich radar BEV features. Extensive experiments on the View-of-Delft (VoD) dataset demonstrate our approach’s superiority over existing methods, achieving the highest mAP of 71.76% across the entire area and 86.36% within the driving corridor. Especially for cars, we improve the AP by 4.17% and 4.20% due to the strong indicative features and symmetric shapes.
Xiangyuan Peng, Huawei Sun, Kay Bierzynski, Anton Fischbacher, Lorenzo Servadei, Robert Wille
ICASSP1
2025 LiRCDepth: Lightweight Radar-Camera Depth Estimation via Knowledge Distillation and Uncertainty Guidance
abstract
Recently, radar-camera fusion algorithms have gained significant attention as radar sensors provide geometric information that complements the limitations of cameras. However, most existing radar-camera depth estimation algorithms focus solely on improving performance, often neglecting computational efficiency. To address this gap, we propose LiRCDepth, a lightweight radar-camera depth estimation model. We incorporate knowledge distillation to enhance the training process, transferring critical information from a complex teacher model to our lightweight student model in three key domains. Firstly, low-level and high-level features are transferred by incorporating pixel-wise and pair-wise distillation. Additionally, we introduce an uncertainty-aware inter-depth distillation loss to refine intermediate depth maps during decoding. Leveraging our proposed knowledge distillation scheme, the lightweight model achieves a 6.6% improvement in MAE on the nuScenes dataset compared to the model trained without distillation. Code: https://github.com/harborsarah/LiRCDepth
Huawei Sun, Nastassia Vysotskaya, Tobias Sukianto, Julius Ott, Xiangyuan Peng, Lorenzo Servadei, Robert Wille
ICASSP6
2025 ELMAR: Enhancing LiDAR Detection with 4D Radar Motion Awareness and Cross-modal Uncertainty
abstract
LiDAR and 4D radar are widely used in autonomous driving and robotics. While LiDAR provides rich spatial information, 4D radar offers velocity measurement and remains robust under adverse conditions. As a result, increasing studies have focused on the 4D radar-LiDAR fusion method to enhance the perception. However, the misalignment between different modalities is often overlooked. To address this challenge and leverage the strengths of both modalities, we propose a LiDAR detection framework enhanced by 4D radar motion status and cross-modal uncertainty. The object movement information from 4D radar is first captured using a Dynamic Motion-Aware Encoding module during feature extraction to enhance 4D radar predictions. Subsequently, the instance-wise uncertainties of bounding boxes are estimated to mitigate the cross-modal misalignment and refine the final LiDAR predictions. Extensive experiments on the View-of-Delft (VoD) dataset highlight the effectiveness of our method, achieving state-of-the-art performance with the mAP of 74.89% in the entire area and 88.70% within the driving corridor while maintaining a real-time inference speed of 30.02 FPS.
Xiangyuan Peng, Huawei Sun, Kay Bierzynski, Lorenzo Servadei, Robert Wille
IROS1
2025 Vector Quantization in the Brain: Grid-like Codes in World Models
abstract
We propose Grid-like Code Quantization (GCQ), a brain-inspired method for compressing observation-action sequences into discrete representations using grid-like patterns in attractor dynamics. Unlike conventional vector quantization approaches that operate on static inputs, GCQ performs spatiotemporal compression through an action-conditioned codebook, where codewords are derived from continuous attractor neural networks and dynamically selected based on actions. This enables GCQ to jointly compress space and time, serving as a unified world model. The resulting representation supports long-horizon prediction, goal-directed planning, and inverse modeling. Experiments across diverse tasks demonstrate GCQ's effectiveness in compact encoding and downstream performance. Our work offers both a computational tool for efficient sequence modeling and a theoretical perspective on the formation of grid-like codes in neural systems.
Xiangyuan Peng, Xingsi Dong, Si Wu 0001
NeurIPS1
2024 MUFASA: Multi-view Fusion and Adaptation Network with Spatial Awareness for Radar Object Detection
Xiangyuan Peng, Huawei Sun, Kay Bierzynski, Lorenzo Servadei, Robert Wille
ICANN (2)1
2024 Visible light communication and WiFi hybrid networks based on dynamic resource allocation algorithm
Boyu Jia, Xue Liang, Xiangyuan Peng
J. Supercomput.5