Xianfei Li

dblp:377/9644 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0001-6641-5212ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 33% Autonomous driving · 16% Vision and language · 9%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
1.722025
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction · ICCV 2025
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction · CVPR 2025
Robotics › Autonomous driving
perception
1.332025
CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous Driving · ACM Multimedia 2024
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction · ICCV 2025
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction · CVPR 2025
Computer vision › 3D vision
3d object detection
0.912025
Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features · ICRA 2025
Computer vision › 3D vision
3d scene understanding
0.912025
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction · CVPR 2025
Computer vision › Vision and language › 3d vision and language
3d vision-language pre-training
0.912025
Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous Driving · AAAI 2025
Machine learning › Deep learning architectures and training › data augmentation
automatic data augmentation
0.912025
Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features · ICRA 2025
Computer vision › 3D vision › 3d object detection
bird's-eye-view detection
0.912025
Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features · ICRA 2025
Machine learning › Deep learning architectures and training
data augmentation
0.912025
Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features · ICRA 2025
Machine learning › Transfer learning and domain adaptation
domain generalization
0.912025
Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features · ICRA 2025
Robotics › Autonomous driving › planning for self-driving vehicles
end-to-end planning
0.912025
Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous Driving · AAAI 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan generation
generative planning
0.912025
Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous Driving · AAAI 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
personalization
0.912025
Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles · ACL (1) 2025
Computer vision › Segmentation and scene understanding › scene understanding
semantic scene understanding
0.912025
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction · ICCV 2025
Machine learning › Transfer learning and domain adaptation › domain generalization
single domain generalization
0.912025
Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features · ICRA 2025
Computer vision › Video understanding and tracking › temporal modeling
temporal fusion
0.912025
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction · CVPR 2025
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
user simulation
0.912025
Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles · ACL (1) 2025
Computer vision › Vision and language
vision-language pretraining
0.912025
Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous Driving · AAAI 2025
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception
0.812024
CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous Driving · ACM Multimedia 2024
Computer vision › 3D vision
camera calibration
0.812024
CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous Driving · ACM Multimedia 2024
Computer vision › 3D vision › camera calibration
multi-camera calibration
0.812024
CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous Driving · ACM Multimedia 2024
Computer vision › 3D vision › 3d scene understanding › semantic scene completion
vision-based occupancy prediction
0.312025
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction · CVPR 2025

Methods — techniques the papers use, named apart from their topics

temporal fusion · 0.9supervised fine-tuning · 0.9reinforcement learning · 0.9recurrent neural network · 0.9gradient descent view · 0.9cycle consistency · 0.9cross-modal language model · 0.9channel-grouped lifting · 0.9causal loss · 0.9autoregressive generation · 0.9
YearPublicationVenuePosition
2026 Improving Proactive Risk-Awareness of Autonomous Driving via Trajectory Monitoring
Xianfei Li, Nanyang Ye 0001
ICPR (4)2
2026 You only click once: single point weakly supervised 3D instance segmentation for autonomous driving
Guangfeng Jiang, Jun Liu 0004, Yongxuan Lv, Yuzhi Wu, Xianfei Li, Wenlong Liao
Expert Syst. Appl.5
2025 Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous Driving
abstract
Autonomous driving is a challenging task that requires perceiving and understanding the surrounding environment for safe trajectory planning. While existing vision-based end-to-end models have achieved promising results, these methods are still facing the challenges of vision understanding, decision reasoning and scene generalization. To solve these issues, a generative planning with 3D-vision language pre-training model named GPVL is proposed for end-to-end autonomous driving. The proposed paradigm has two significant aspects. On one hand, a 3D-vision language pre-training module is designed to bridge the gap between visual perception and linguistic understanding in the bird's eye view. On the other hand, a cross-modal language model is introduced to generate reasonable planning with perception and navigation information in an auto-regressive manner. Experiments on the challenging nuScenes dataset demonstrate that the proposed scheme achieves excellent performances compared with state-of-the-art methods. Besides, the proposed GPVL presents strong generalization ability and real-time potential when handling high-level commands in various scenarios. It is believed that the effective, robust and efficient performance of GPVL is crucial for the practical application of future autonomous driving systems.
Tengpeng Li, Hanli Wang, Xianfei Li, Wenlong Liao
AAAI3
2025 Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles
abstract
User simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and user-level diversity, often hindered by role confusion and dependence on predefined profiles of well-known figures. In contrast, direct simulation focuses solely on text, neglecting implicit user traits like personality and conversation-level consistency. To address these issues, we introduce the User Simulator with Implicit Profiles (USP), a framework that infers implicit user profiles from human-machine interactions to simulate personalized and realistic dialogues. We first develop an LLM-driven extractor with a comprehensive profile schema, then refine the simulation using conditional supervised fine-tuning and reinforcement learning with cycle consistency, optimizing at both the utterance and conversation levels. Finally, a diverse profile sampler captures the distribution of real-world user profiles. Experimental results show that USP outperforms strong baselines in terms of authenticity and diversity while maintaining comparable consistency. Additionally, using USP to evaluate LLM on dynamic multi-turn aligns well with mainstream benchmarks, demonstrating its effectiveness in real-world applications.
Kuang Wang, Xianfei Li, Shenghao Yang 0001, Li Zhou 0010, Feng Jiang 0007, Haizhou Li 0001
ACL (1)2
2025 Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
abstract
We present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and fusion strategies. It systematically examines the entire VisionOcc pipeline, identifying three fundamental yet previously overlooked temporal cues: scene-level consistency, motion calibration, and geometric complementation. These cues capture diverse facets of temporal evolution and make distinct contributions across various modules in the VisionOcc framework. To effectively fuse temporal signals across heterogeneous representations, we propose a novel fusion strategy by reinterpreting the formulation of vanilla RNNs. This reinterpretation leverages gradient descent on features to unify the integration of diverse temporal information, seamlessly embedding the proposed temporal cues into the network. Extensive experiments on nuScenes demonstrate that GDFusion significantly outperforms established baselines, achieving 2.2%–4.7% mIoU improvement and reducing memory consumption by 30%–72%. Codes are available at https: //github.com/cdb342/GDFusion.
Dubing Chen, Xingping Dong, Xianfei Li, Wenlong Liao, Jianbing Shen
CVPR5
2025 Semantic Causality-Aware Vision-Based 3D Occupancy Prediction
abstract
Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leading to cascading errors. In this paper, we address this limitation by designing a novel causal loss that enables holistic, end-to-end supervision of the modular 2D-to-3D transformation pipeline. Grounded in the principle of 2D-to-3D semantic causality, this loss regulates the gradient flow from 3D voxel representations back to the 2D features. Consequently, it renders the entire pipeline differentiable, unifying the learning process and making previously non-trainable components fully learnable. Building on this principle, we propose the Semantic Causality-Aware 2D-to-3D Transformation, which comprises three components guided by our causal loss: Channel-Grouped Lifting for adaptive semantic mapping, Learnable Camera Offsets for enhanced robustness against camera perturbations, and Normalized Convolution for effective feature propagation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Occ3D benchmark, demonstrating significant robustness to camera perturbations and improved 2D-to-3D semantic consistency.
Dubing Chen, Yucheng Zhou 0001, Xianfei Li, Wenlong Liao, Jianbing Shen
ICCV4
2025 Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D Features
abstract
With the increasing popularity of autonomous driving based on the Bird's-Eye-View (BEV) representation, improving the generalization of such detection models is key for safe real-world applications. However, a realistic yet challenging scenario: Single Domain Generalization (SDG) for BEV, is still under-explored. A key ingredient for SDG is to increase data diversity via common image augmentation or adversarial data generation first. However, common image-level augmentation is not sufficient enough to ensure domain diversity in most part of latent space. The adversarial generation has the problem of unstable training or mode collapsing as well. To address these limitations, we present Tri-level Automatic Augmentation (Tri-AutoAug), a simple yet effective method to enlarge the diversity and quantity of data from image and 2D features and facilitate the model to learn more domain-invariant features in BEV space. Besides, Tri-AutoAug can automatically learn augmentation strategies to avoid spending too much time manually adjusting hyperparameters and maximize the benefit of Tri-level Augmentation. To the best of our knowledge, this is the first study to explore automatic augmentation for SDG BEV. Extensive experiments on NuScenes-C including eight testing domains have demonstrated that our approach can achieve the best performance across various domain generalization methods. More importantly, we evaluate the proposed method in real-world autonomous driving scenarios. Tri-AutoAug improves the out-of-distribution (ood) performance by 8.54% (mAP), which demonstrates that Tri-AutoAug provides a practical and feasible solution for the applications of 3D detectors in the real world. The code is available at https://github.com/ClaireTunlTri-AutoAug.
Xianfei Li, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001
ICRA3
2025 Open-vocabulary object detection via Neighboring Region Attention Alignment
Sunyuan Qiang, Xianfei Li, Yanyan Liang 0001, Wenlong Liao
Eng. Appl. Artif. Intell.2
2024 CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous Driving
abstract
Camera calibration is crucial in computer vision tasks and applications, e.g., autonomous driving (AD). However, prevailing camera calibration models pose a time-consuming and labor-intensive off-board process in mass production settings, while simultaneously lacking exploration of real-world AD scenarios. To this end, inspired by recent advancements in bird's-eye-view (BEV) perception models, this paper proposes a novel multi-camera Calibration method via Reversed BEV representations for AD, termed CalibRBEV. Specifically, the proposed CalibRBEV model primarily comprises two stages. Initially, we innovatively reverse the BEV perception pipeline, reconstructing bounding boxes through an attention auto-encoder module to fully extract the latent reversed BEV representations. Subsequently, the obtained representations from encoder are interacted with the surrounding multi-view image features for further refinement and calibration parameters prediction. Extensive experimental results on nuScenes and Waymo datasets validate the effectiveness of our proposed model.
Wenlong Liao, Sunyuan Qiang, Xianfei Li, Yanyan Liang 0001, Junchi Yan
ACM Multimedia3