Xingcheng Zhou

dblp:317/7174 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 28% Video understanding and tracking · 22% Autonomous driving · 20%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
end-to-end driving
1.012026
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model · AAAI 2026
Robotics › Motion planning and robot control
trajectory planning
1.012026
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model · AAAI 2026
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.012026
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model · AAAI 2026
Computer vision › Vision and language › image captioning › grounded image captioning
object captioning
0.912025
TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes · ICML 2025
Computer vision › Vision and language › visual grounding
object grounding
0.912025
TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes · ICML 2025
Computer vision › Video understanding and tracking › spatio-temporal understanding
spatio-temporal video understanding
0.912025
TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes · ICML 2025
Computer vision › Video understanding and tracking
video question answering
0.912025
TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes · ICML 2025
Robotics › Autonomous driving
collaborative perception
0.812024
TUMTraf V2X Cooperative Perception Dataset · CVPR 2024
Computer vision › 3D vision › 3d object detection
multi-agent 3d object detection
0.812024
TUMTraf V2X Cooperative Perception Dataset · CVPR 2024
Computer vision › Vision and language
multimodal fusion
0.812024
TUMTraf V2X Cooperative Perception Dataset · CVPR 2024
Computer vision › Video understanding and tracking
multi-object tracking
0.212024
TUMTraf V2X Cooperative Perception Dataset · CVPR 2024

Methods — techniques the papers use, named apart from their topics

vision-language alignment · 1.0large language model · 1.0autoregressive decoding · 1.0visual token sampling · 0.9multiple-choice QA · 0.9multimodal baseline · 0.9coopdet3d · 0.8camera-LiDAR fusion · 0.8
YearPublicationVenuePosition
2026 OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
abstract
We present OpenDriveVLA, a Vision-Language Action (VLA) model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially-grounded driving actions by leveraging multimodal inputs, including both 2D and 3D instance-aware visual representations, ego vehicle states, and language commands. To bridge the modality gap between driving visual representations and language embeddings, we introduce a hierarchical vision-language alignment process, projecting both 2D and 3D structured visual tokens into a unified semantic space. Furthermore, we incorporate structured agent–environment–ego interaction modeling into the autoregressive decoding process, enabling the model to capture fine-grained spatial dependencies and behavior-aware dynamics critical for reliable trajectory planning. Extensive experiments on the nuScenes dataset demonstrate that OpenDriveVLA achieves state-of-the-art results across open-loop trajectory planning and driving-related question-answering tasks. Qualitative analyses further illustrate its superior capability to follow high-level driving commands and robustly generate trajectories under challenging scenarios, highlighting its potential for next-generation end-to-end autonomous driving.
Xingcheng Zhou, Xuyuan Han, Yunpu Ma, Volker Tresp, Alois C. Knoll
AAAI1
2025 TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
abstract
We present TUMTraf VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraf VideoQA unifies three essential tasks—multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding—within a cohesive evaluation framework. We further introduce the TraffiX-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset’s complexity, highlight the limitations of existing models, and position TUMTraf VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.
Xingcheng Zhou, Konstantinos Larintzakis, Walter Zimmer, Hu Cao, Venkatnarayanan Lakshminarasimhan, Leah Strand, Alois C. Knoll
ICML1
2025 MambaSFLNet: A Mamba-based Model for Low-Light Image Enhancement with Spatial and Frequency Features
abstract
Low-light image enhancement (LLIE) aims to enhance the illumination of images that are captured under dark conditions, which is critical for various applications in dim environments, such as robotics and autonomous driving. Existing convolutional neural network (CNN)-based methods usually struggle to capture long-range dependencies, while transformer-based methods, despite their effectiveness, are resource-consuming. Besides, the frequency domain includes important lightness degradation information. To this end, we propose a Mamba-based framework called MambaSFLNet to effectively address LLIE by integrating spatial and frequency features. Our approach utilizes the Visual State Space Module to establish relationships across different regions of the input image while maintaining low model complexity. Furthermore, The spatial module not only balances illumination distribution but also suppresses noise and artifacts during enhancement. In addition, the frequency module enhances image contrast and sharpness by leveraging frequency-domain information. Extensive experiments on nine widely used benchmarks demonstrate that our approach achieves superior performance and exhibits strong generalization capabilities compared to existing methods. The codes are available at https://github.com/MingyuLiu1/MambaSFLNet.git
Yuning Cui 0001, Leah Strand, Xingcheng Zhou, Alois C. Knoll
IROS4
2025 LiDAR-Guided Monocular 3D Object Detection for Long-Range Railway Monitoring
abstract
Railway systems, particularly in Germany, require high levels of automation to address legacy infrastructure challenges and increase train traffic safely. A key component of automation is robust long-range perception, essential for early hazard detection, such as obstacles at level crossings or pedestrians on tracks. Unlike automotive systems with braking distances of 70 meters, trains require perception ranges exceeding 1 km. This paper presents an deep-learning-based approach for long-range 3D object detection tailored for autonomous trains. The method relies solely on monocular images, inspired by the Faraway-Frustum approach, and incorporates LiDAR data during training to improve depth estimation. The proposed pipeline consists of four key modules: (1) a modified YOLOv9 for 2.5D object detection, (2) a depth estimation network, and (3–4) dedicated short- and long-range 3D detection heads. Evaluations on the OSDaR23 dataset demonstrate the effectiveness of the approach in detecting objects up to 250 meters. Results highlight its potential for railway automation and outline areas for future improvement.
Raul David Dominguez Sanchez, Xavier Jair Diaz Ortiz, Xingcheng Zhou, Max Peter Ronecker, Michael Karner, Daniel Watzenig, Alois C. Knoll
IV3
2024 TUMTraf V2X Cooperative Perception Dataset
abstract
Cooperative perception offers several benefits for en-hancing the capabilities of autonomous vehicles and im-proving road safety. Using roadside sensors in addition to onboard sensors increases reliability and extends the sensor range. External sensors offer higher situational awareness for automated vehicles and prevent occlusions. We propose CoopDet3D, a cooperative multi-modal fusion model, and TUMTraf- V2X, a perception dataset, for the cooperative 3D object detection and tracking task. Our dataset contains 2,000 labeled point clouds and 5,000 labeled images from five roadside and four onboard sensors. It includes 30k 3D boxes with track IDs and precise GPS and IMU data. We labeled nine categories and covered occlusion scenarios with challenging driving maneuvers, like traffic violations, near-miss events, overtaking, and U-turns. Through multiple experiments, we show that our CoopDet3D camera-LiDARfusion model achieves an increase of +14.36 3D mAP compared to a vehicle camera-LiDARfusion model. Finally, we make our dataset, model, labeling tool, and devkit publicly available on our website.
Walter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou, Rui Song 0007, Alois C. Knoll
CVPR4