VLDB 2026 Research / reviewers in the wild / expert
Rui Song 0007
dblp:01/2743-7
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-7359-1081ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Drivingabstract28031 Rui Song 0007, Chenwei Liang, Yan Xia 0003, Walter Zimmer, Hu Cao, Holger Caesar, Andreas Festag, Alois C. Knoll |
ICCV | 1 |
| 2025 | Feature-aligned Fisheye Object Detection Network for Autonomous DrivingabstractFisheye cameras, renowned for their panoramic field of view (FOV) of 360°, are crucial for surround-view perception in autonomous driving. However, research on object perception in fisheye images lags behind that of standard images. To address this gap, we propose a feature-aligned fisheye object detection network specifically tailored for autonomous driving. Current fisheye perception algorithms often overlook the misalignment issues that typically arise in object detectors. To tackle these challenges in the feature pyramid network (FPN), we introduce a feature-aligned pyramid module (FaPM), which learns pixel transformation offsets to contextually align feature maps. Additionally, we present a location-aligned detection head (LaDH) to align the spatial distribution of classification and regression localization. Integrating these modules into a detection framework results in a novel feature-aligned fisheye object detector. Our method undergoes extensive evaluation on the WoodScape dataset, achieving a mean average precision (mAP) of 32.2%, surpassing the performance of existing methods. Hu Cao, Dongyi Sun, Rui Song 0007, Yan Xia 0003, Alois C. Knoll |
IROS | 3 |
| 2025 | V2X-Gaussians: Gaussian Splatting for Multi-Agent Cooperative Dynamic Scene ReconstructionabstractRecent advances in neural rendering, such as NeRF and Gaussian Splatting, have shown great potential for dynamic scene reconstruction in intelligent vehicles. However, existing methods rely on a single ego-vehicle, suffering from limited field-of-view and occlusions, leading to incomplete reconstructions. While V2X communication may provide additional information from roadside infrastructure or other vehicless, it often degrades reconstruction quality due to sparse overlapping views. In this paper, we propose V2X-Gaussians, the first framework integrating V2X communication into Gaussian Splatting. Specifically, by leveraging deformable Gaussians and an iterative V2X-aware cross-ray densification approach, we enhance infrastructure-aided neural rendering and address view sparsity in multi-agent cooperative scenarios. In addition, to support systematic evaluation, we introduce a standardized benchmark for V2X scene reconstruction. Experiments on real-world data show that our method outperforms state-of-the-art approaches by +2.09 PSNR with only 561.8 KB for periodic V2X data exchange, highlighting the benefits of incorporating roadside infrastructure into neural rendering for intelligent transportation systems. Our code and benchmark are publicly available under an open-source license.41https://github.com/abhishekjagtap1/V2X-Guassians Abhishek Dinkar Jagtap, Rui Song 0007, Sanath Tiptur Sadashivaiah, Andreas Festag |
IV | 2 |
| 2025 | Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localizationabstractAutonomous vehicles often rely on high-definition (HD) maps for navigation; however, these maps are not frequently updated and often lack semi-static information, such as temporary roadwork zones, which can significantly alter the road network. This limitation underscores the urgent need for an accurate global position of roadwork zones. However, the absence of publicly available datasets for evaluating roadwork zone detection and geo-localization models has hindered the development of reliable autonomous driving systems. To address this challenge, we propose the Roadwork Zone Detection and Geo-localization (RZDG) dataset, which includes both simulated and real-world data, providing multimodal sensor inputs along with comprehensive annotations. The dataset supports multiple perception tasks, including image semantic segmentation, 3D object detection, and object geo-localization. In addition, we introduce a tracker-based roadwork zone detection and geo-localization (RZDG) pipeline, an extension of AB3DMOT, for accurate object geo-localization in roadwork zones. We benchmark our approach on the RZDG dataset, demonstrating its effectiveness in detecting roadwork zones and transforming object positions from the local coordinate system to the global coordinate system. A prediction is considered a true positive (TP) if its estimated position falls within one meter of the ground truth. Our experimental results show that our approach achieves high accuracy on both real and simulated data. Specifically, we report: Precision: 0.565 (real) / 0.615 (simulated) Recall: 0.898 (real) / 0.809 (simulated) F1-score: 0.597 (real) / 0.665 (simulated). The RZDG dataset and code can be found at: https://github.com/chrisyan/RZDG. Zhiran Yan, Yutong Xin, S. Shyam Shenoi, Rui Song 0007, Gordon Elger |
IV | 4 |
| 2024 | Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated VehiclesabstractCollaborative perception in automated vehicles lever-ages the exchange of information between agents, aiming to elevate perception results. Previous camera-based collabo-rative 3D perception methods typically employ 3D bounding boxes or bird's eye views as representations of the en-vironment. However, these approaches fall short in offering a comprehensive 3D environmental prediction. To bridge this gap, we introduce the first method for collaborative 3D semantic occupancy prediction. Particularly, it improves local 3D semantic occupancy predictions by hybrid fusion of (i) semantic and occupancy task features, and (ii) Compressed orthogonal attention features shared between vehi-cles. Additionally, due to the lack of a collaborative perception dataset designed for semantic occupancy prediction, we augment a current collaborative perception dataset to include 3D collaborative semantic occupancy labels for a more robust evaluation. The experimental findings highlight that: (i) our collaborative semantic occupancy predictions excel above the results from single vehicles by over 30%, and (ii) models anchored on semantic occupancy outpace state-of-the-art collaborative 3D detection techniques in subsequent perception applications, showcasing enhanced accuracy and enriched semantic-awareness in road environments. Rui Song 0007, Chenwei Liang, Hu Cao, Zhiran Yan, Walter Zimmer, Markus Gross 0003, Andreas Festag, Alois C. Knoll |
CVPR | 1 |
| 2024 | TUMTraf V2X Cooperative Perception DatasetabstractCooperative perception offers several benefits for en-hancing the capabilities of autonomous vehicles and im-proving road safety. Using roadside sensors in addition to onboard sensors increases reliability and extends the sensor range. External sensors offer higher situational awareness for automated vehicles and prevent occlusions. We propose CoopDet3D, a cooperative multi-modal fusion model, and TUMTraf- V2X, a perception dataset, for the cooperative 3D object detection and tracking task. Our dataset contains 2,000 labeled point clouds and 5,000 labeled images from five roadside and four onboard sensors. It includes 30k 3D boxes with track IDs and precise GPS and IMU data. We labeled nine categories and covered occlusion scenarios with challenging driving maneuvers, like traffic violations, near-miss events, overtaking, and U-turns. Through multiple experiments, we show that our CoopDet3D camera-LiDARfusion model achieves an increase of +14.36 3D mAP compared to a vehicle camera-LiDARfusion model. Finally, we make our dataset, model, labeling tool, and devkit publicly available on our website. Walter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou, Rui Song 0007, Alois C. Knoll |
CVPR | 5 |
| 2024 | First Mile: An Open Innovation Lab for Infrastructure-Assisted Cooperative Intelligent Transportation SystemsabstractInfrastructure-assisted Cooperative Intelligent Transportation Systems (C-ITS) leverage roadside intelligent infrastructure and vehicular network technology to facilitate information exchange among traffic participants, enhancing road safety, efficiency, and sustainability. However, this requires not only the massive deployment of infrastructure, including advanced sensors, communication devices, and computing units at various levels, but also the involvement of various stakeholders in the development of functions and real-road testing. In this paper, we present our test field – First Mile, as an open innovation lab for C-ITS in Ingolstadt, Germany, offering an open environment for world-wide stakeholders to conduct research and testing in C-ITS. In particular, we equip a 3.5 km area with 22 roadside intelligent masts and 89 sensors, achieving dense deployment on public roads. Our design includes a protocol stack tailored for different C-ITS stations and services, conforming to European communication standards. Furthermore, we conduct quantitative analyzes of key performance metrics, such as radio signal quality and End-to-End delay, to assess the efficacy of First Mile in different data processing pipelines. Finally, we delve into the future prospects of large-scale C-ITS deployment, guided by extensive and prolonged measurement studies. Rui Song 0007, Andreas Festag, Abhishek Dinkar Jagtap, Maximilian Bialdyga, Zhiran Yan, Maximilian Otte, Sanath Tiptur Sadashivaiah, Alois C. Knoll |
IV | 1 |
| 2024 | ResFed: Communication0Efficient Federated Learning With Deep Compressed ResidualsabstractFederated learning allows for cooperative training among distributed clients by sharing their locally learned model parameters, such as weights or gradients. However, as model size increases, the communication bandwidth required for deployment in wireless networks becomes a bottleneck. To address this, we propose a residual-based federated learning framework (ResFed) that transmits residuals instead of gradients or weights in networks. By predicting model updates at both clients and the server, residuals are calculated as the difference between updated and predicted models and contain more dense information than weights or gradients. We find that the residuals are less sensitive to an increasing compression ratio than other parameters, and hence use lossy compression techniques on residuals to improve communication efficiency for training in federated settings. With the same compression ratio, ResFed outperforms current methods (weight-or gradient-based federated learning) by over 1.4× on federated datasets, including MNIST, FashionMNIST, SVHN, CIFAR-10, CIFAR-100, FEMNIST, in client-to-server communication, and can also be applied to reduce communication costs for server-to-client communication. Rui Song 0007, Liguo Zhou, Lingjuan Lyu, Andreas Festag, Alois C. Knoll |
IEEE Internet Things J. | 1 |
| 2023 | V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative PerceptionabstractModern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at research.seas.ucla.edu/mobility-lab/v2v4real/. Runsheng Xu, Xin Xia 0007, Hanzhao Li, Zhengzhong Tu, Zonglin Meng, Hao Xiang 0001, Rui Song 0007, Hongkai Yu, Bolei Zhou, Jiaqi Ma 0003 |
CVPR | 10 |
| 2023 | Federated Learning via Decentralized Dataset Distillation in Resource-Constrained Edge EnvironmentsabstractIn federated learning, all networked clients contribute to the model training cooperatively. However, with model sizes increasing, even sharing the trained partial models often leads to severe communication bottlenecks in underlying networks, especially when communicated iteratively. In this paper, we introduce a federated learning framework FedD3 requiring only one-shot communication by integrating dataset distillation instances. Instead of sharing model updates in other federated learning approaches, FedD3 allows the connected clients to distill the local datasets independently, and then aggregates those decentralized distilled datasets (e.g. a few unrecognizable images) from networks for model training. Our experimental results show that FedD3 significantly outperforms other federated learning frameworks in terms of needed communication volumes, while it provides the additional benefit to be able to balance the trade-off between accuracy and communication cost, depending on usage scenario or target dataset. For instance, for training an AlexNet model on CIFAR-10 with 10 clients under non-independent and identically distributed (Non-IID) setting, FedD3 can either increase the accuracy by over 71% with a similar communication volume, or save 98% of communication volume, while reaching the same accuracy, compared to other one-shot federated learning approaches. Rui Song 0007, Dai Liu, Dave Zhenyu Chen, Andreas Festag, Carsten Trinitis, Martin Schulz 0001, Alois C. Knoll |
IJCNN | 1 |
| 2023 | Residual encoding framework to compress DNN parameters for fast transfer
Liguo Zhou, Rui Song 0007, Guang Chen 0001, Andreas Festag, Alois C. Knoll |
Knowl. Based Syst. | 2 |
| 2022 | Concept of Smart Infrastructure for Connected Vehicle Assist and Traffic Flow Optimizationabstract360 Shiva Agrawal, Rui Song 0007, Akhil Kohli, Andreas Korb, Maximilian Andre, Erik Holzinger, Gordon Elger |
VEHITS | 2 |
| 2022 | Edge-Aided Sensor Data Sharing in Vehicular Communication NetworksabstractSensor data sharing in vehicular networks can significantly improve the range and accuracy of environmental perception for connected automated vehicles. Different concepts and schemes for dissemination and fusion of sensor data have been developed. It is common to these schemes that measurement errors of the sensors impair the perception quality and can result in road traffic accidents. Specifically, when the measurement error from the sensors - also referred as measurement noise - is unknown and time varying, the performance of the data fusion process is restricted, which represents a major challenge in the calibration of sensors. In this paper, we consider sensor data sharing and fusion in a vehicular network with both, vehicle-to-infrastructure and vehicle-to-vehicle communication. We propose a method, named Bidirectional Feedback Noise Estimation (BiFNoE), in which an edge server collects and caches sensor measurement data from vehicles. The edge estimates the noise and the targets alternately in double dynamic sliding time windows and enhances the distributed cooperative environment sensing at each vehicle with low communication costs. We evaluate the proposed algorithm and data dissemination strategy in an application scenario by simulation and show that the perception accuracy is on average improved by around 80% with only 12 kbps uplink and 28 kbps downlink bandwidth. Rui Song 0007, Anupama Hegde, Numan Senel, Alois C. Knoll, Andreas Festag |
VTC Spring | 1 |