EDBT 2026 Demo / reviewers in the wild / expert
Qingwen Zhang
dblp:83/6437
· DBLP profile ↗
12ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint)abstractLiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for non-ego motion compensation, correcting the representation of dynamic objects in point clouds. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles. Qingwen Zhang, Ajinkya Khoche, Yi Yang 0095, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt |
AAAI | 1 |
| 2025 | AGO: Adaptive Grounding for Open World 3D Occupancy PredictionabstractOpen-world 3D semantic occupancy prediction aims to generate a voxelized 3D representation from sensor inputs while recognizing both known and unknown objects. Transferring open-vocabulary knowledge from vision-language models (VLMs) offers a promising direction but remains challenging. However, methods based on VLM-derived 2D pseudo-labels with traditional supervision are limited by a predefined label space and lack general prediction capabilities. Direct alignment with pretrained image embeddings, on the other hand, often fails to achieve reliable performance because of inconsistent image and text representations in VLMs. To address these challenges, we propose AGO, a novel 3D occupancy prediction framework with adaptive grounding to handle diverse open-world scenarios. AGO first encodes surrounding images and class prompts into 3D and text embeddings, respectively, leveraging similarity-based grounding training with 3D pseudo-labels. Additionally, a modality adapter maps 3D embeddings into a space aligned with VLM-derived image embeddings, reducing modality gaps. Experiments on Occ3D-nuScenes show that AGO improves unknown object prediction in zero-shot and few-shot transfer while achieving state-of-the-art closed-world self-supervised performance, surpassing prior methods by 4.09 mIoU. Code is available at: https://github.com/EdwardLeeLPZ/AGO. Peizheng Li, Shuxiao Ding, Qingwen Zhang, Onat Inak, Larissa Triess, Niklas Hanselmann, Marius Cordts, Andreas Zell |
ICCV | 4 |
| 2025 | SSF: Sparse Long-Range Scene Flow for Autonomous DrivingabstractScene flow enables an understanding of the motion characteristics of the environment in the 3D world. It gains particular significance in the long-range, where object-based perception methods might fail due to sparse observations far away. Although significant advancements have been made in scene flow pipelines to handle large-scale point clouds, a gap remains in scalability with respect to long-range. We attribute this limitation to the common design choice of using dense feature grids, which scale quadratically with range. In this paper, we propose Sparse Scene Flow (SSF), a general pipeline for long-range scene flow, adopting a sparse convolution based backbone for feature extraction. This approach introduces a new challenge: a mismatch in size and ordering of sparse feature maps between time-sequential point scans. To address this, we propose a sparse feature fusion scheme, that augments the feature maps with virtual voxels at missing locations. Additionally, we propose a range-wise metric that implicitly gives greater importance to faraway points. Our method, SSF, achieves state-of-the-art results on the Argoverse2 dataset, demonstrating strong performance in long-range scene flow estimation. Our code is open-sourced at https://github.com/KTH-RPL/SSF.git. Ajinkya Khoche, Qingwen Zhang, Laura Pereira Sánchez, Aron Asefaw, Sina Sharif Mansouri, Patric Jensfelt |
ICRA | 2 |
| 2025 | DeltaFlow: An Efficient Multi-frame Scene Flow Estimation MethodabstractPrevious dominant methods for scene flow estimation focus mainly on input from two consecutive frames, neglecting valuable information in the temporal domain. While recent trends shift towards multi-frame reasoning, they suffer from rapidly escalating computational costs as the number of frames grows. To leverage temporal information more efficiently, we propose DeltaFlow ($\Delta$Flow), a lightweight 3D framework that captures motion cues via a $\Delta$ scheme, extracting temporal features with minimal computational cost, regardless of the number of frames. Additionally, scene flow estimation faces challenges such as imbalanced object class distributions and motion inconsistency. To tackle these issues, we introduce a Category-Balanced Loss to enhance learning across underrepresented classes and an Instance Consistency Loss to enforce coherent object motion, improving flow accuracy. Extensive evaluations on the Argoverse 2, Waymo and nuScenes datasets show that $\Delta$Flow achieves state-of-the-art performance with up to 22\% lower error and $2\times$ faster inference compared to the next-best multi-frame supervised method, while also demonstrating a strong cross-domain generalization ability. The code is open-sourced at https://github.com/Kin-Zhang/DeltaFlow along with trained model weights. Qingwen Zhang, Yushan Zhang, Yixi Cai, Olov Andersson, Patric Jensfelt |
NeurIPS | 1 |
| 2025 | HiMo: High-Speed Objects Motion Compensation in Point CloudsabstractLiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for non-ego motion compensation, correcting the representation of dynamic objects in point clouds. During the development of HiMo, we observed that existing self-supervised scene flow estimators often produce degenerate or inconsistent estimates under high-speed distortion. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. Since well-established motion distortion metrics are absent in the literature, we introduce two evaluation metrics: compensation accuracy at a point level and shape similarity of objects. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles. Our findings show that HiMo improves the geometric consistency and visual fidelity of dynamic objects in LiDAR point clouds, benefiting downstream tasks such as semantic segmentation and 3D detection. See https://kin-zhang.github.io/HiMo for more details. Qingwen Zhang, Ajinkya Khoche, Yi Yang 0095, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt |
IEEE Trans. Robotics | 1 |
| 2024 | SeFlow: A Self-supervised Scene Flow Method in Autonomous Driving
Qingwen Zhang, Yi Yang 0095, Peizheng Li, Olov Andersson, Patric Jensfelt |
ECCV (1) | 1 |
| 2024 | DeFlow: Decoder of Scene Flow Network in Autonomous DrivingabstractScene flow estimation determines a scene’s 3D motion field, by predicting the motion of points in the scene, especially for aiding tasks in autonomous driving. Many networks with large-scale point clouds as input use voxelization to create a pseudo-image for real-time running. However, the voxelization process often results in the loss of point-specific features. This gives rise to a challenge in recovering those features for scene flow tasks. Our paper introduces DeFlow which enables a transition from voxel-based features to point features using Gated Recurrent Unit (GRU) refinement. To further enhance scene flow estimation performance, we formulate a novel loss function that accounts for the data imbalance between static and dynamic points. Evaluations on the Argoverse 2 scene flow task reveal that DeFlow achieves state-of-the-art results on large-scale point cloud data, demonstrating that our network has better performance and efficiency compared to others. The code is available at https://github.com/KTH-RPL/deflow. Qingwen Zhang, Yi Yang 0095, Ruoyu Geng, Patric Jensfelt |
ICRA | 1 |
| 2024 | Hard Cases Detection in Motion Prediction by Vision-Language Foundation ModelsabstractAddressing hard cases in autonomous driving, such as anomalous road users, extreme weather conditions, and complex traffic interactions, presents significant challenges. To ensure safety, it is crucial to detect and manage these scenarios effectively for autonomous driving systems. However, the rarity and high-risk nature of these cases demand extensive, diverse datasets for training robust models. Vision-Language Foundation Models (VLMs) have shown remarkable zero-shot capabilities as being trained on extensive datasets. This work explores the potential of VLMs in detecting hard cases in autonomous driving. We demonstrate the capability of VLMs such as GPT-4v in detecting hard cases in traffic participant motion prediction on both agent and scenario levels. We introduce a feasible pipeline where VLMs, fed with sequential image frames with designed prompts, effectively identify challenging agents or scenarios, which are verified by existing prediction models. Moreover, by taking advantage of this detection of hard cases by VLMs, we further improve the training efficiency of the existing motion prediction pipeline by performing data selection for the training samples suggested by GPT. We show the effectiveness and feasibility of our pipeline incorporating VLMs with state-of-the-art methods on NuScenes datasets. The code is accessible at https://github.com/KTH-RPL/Detect_VLM. Yi Yang 0095, Qingwen Zhang, Kei Ikemura, Nazre Batool, John Folkesson |
IV | 2 |
| 2023 | A VT-HMM-Based Framework for Countdown Timer Traffic Light State EstimationabstractTraffic lights are important components of traffic systems, and perceptual tasks on traffic lights are crucial for intelligent agents on the road. Auxiliary countdown timers, providing the remaining time of the current traffic phase, improve the safety and smoothness of the entire traffic system. This work proposes a state estimation framework for countdown timer traffic lights. Time-domain information is adequately integrated into a variable transition Hidden Markov Model (VT-HMM), and our system provides optimal estimates of traffic light colors and countdown numbers based on noisy detection inputs. A dynamic state transition matrix is designed based on a 1-step transition logic and a probability of the number of transitions related to the current state sojourn duration. A recursive decoding method based on the Viterbi algorithm is proposed to update all the state candidates and select the optimal state chain. Extensive experiments evaluate the robustness and effectiveness of the proposed work. The performance boundaries of this system are also found under various input noise levels. The source code is available here:https://github.com/ShuyangUni/countdown-timer-traffic-light-estimation Qingwen Zhang, Feiyi Chen, Jin Wu 0002, Jianhao Jiao, Lujia Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Real-Time Trajectory Planning for Autonomous Driving with Gaussian Process and Incremental RefinementabstractReal-time kinodynamic trajectory planning in dy-namic environments is critical yet challenging for autonomous driving. In this paper, we propose an efficient trajectory plan-ning system for autonomous driving in complex dynamic sce-narios through iterative and incremental path-speed optimization. Exploiting the decoupled structure of the planning prob-lem, a path planner based on Gaussian process first generates a continuous arc-length parameterized path in the Frenét frame, considering static obstacle avoidance and curvature constraints. We theoretically prove that it is a good generalization of the well-known jerk optimal solution. An efficient s-t graph search method is introduced to find a speed profile along the generated path to deal with dynamic environments. Finally, the path and speed are optimized incrementally and iteratively to ensure kinodynamic feasibility. Various simulated scenarios with both static obstacles and dynamic agents verify the effectiveness and robustness of our proposed method. Experimental results show that our method can run at 20 Hz. The source code is released as an open-source package. Jie Cheng 0008, Yingbing Chen, Qingwen Zhang, Lu Gan 0001, Ming Liu 0001 |
ICRA | 3 |
| 2022 | MMFN: Multi-Modal-Fusion-Net for End-to-End DrivingabstractInspired by the fact that humans use diverse sensory organs to perceive the world, sensors with different modalities are deployed in end-to-end driving to obtain the global context of the 3D scene. In previous works, camera and LiDAR inputs are fused through transformers for better driving performance. These inputs are normally further interpreted as high-level map information to assist navigation tasks. Nevertheless, extracting useful information from the complex map input is challenging, for redundant information may mislead the agent and negatively affect driving performance. We propose a novel approach to efficiently extract features from vectorized High-Definition (HD) maps and utilize them in end-to-end driving tasks. In addition, we design a new expert to enhance the model performance by considering multi-road rules. Experimental results prove that both proposed improvements enable our agent to achieve superior performance compared with other methods. Qingwen Zhang, Mingkai Tang 0002, Ruoyu Geng, Feiyi Chen, Ren Xin, Lujia Wang 0001 |
IROS | 1 |
| 2016 | Simulation of non-point source pollution load in the Xiangtan Stream basin through swat modelabstractXiangtan Stream basin, a typical hilly country in the Sichuan Basin, is adversely affected by overuse of chemical fertilizer on the slope farmland and pig manure produced by intensive pig breeding farm located in the upper watershed. The nitrogen and phosphorus compounds in the runoff, sediment and manure contaminate the soil and water bodies. SWAT (Soil Water and Assessment Tool) was applied to characterize hydrologic processes in the watershed and to evaluate the reduction effect of non-point source pollution (NPSP). Field data including flow and concentration (total phosphorus (TP), total nitrogen (TN) and nitrate-nitrogen (NO3-N)) were measured for the time period from March 2012 to March 2013. The result of comparison showed that SWAT model performed satisfactorily for simulating runoff and water quality in the study area. TN and TP load was found to be concentrated in the upper parts of the watershed, critical areas with the pig breeding farm. Contribution rate analysis suggested that livestock was the major contributor for TN and TP load, which should be the key factor for NSPS. Temporal analysis showed, monthly output of TN and TP was higher in June and July. Huazhang Liu, Shihua Li 0002, Qingwen Zhang |
IGARSS | 3 |