Ajinkya Khoche

dblp:332/1939 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0009-6935-6797ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 58% Motion planning and robot control · 23% Autonomous driving · 19%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
scene flow estimation
2.732026
HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint) · AAAI 2026
HiMo: High-Speed Objects Motion Compensation in Point Clouds · IEEE Trans. Robotics 2025
SSF: Sparse Long-Range Scene Flow for Autonomous Driving · ICRA 2025
Robotics › Motion planning and robot control › robot control
motion compensation
1.922026
HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint) · AAAI 2026
HiMo: High-Speed Objects Motion Compensation in Point Clouds · IEEE Trans. Robotics 2025
Robotics › Autonomous driving › perception
LiDAR perception
1.322026
HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint) · AAAI 2026
HiMo: High-Speed Objects Motion Compensation in Point Clouds · IEEE Trans. Robotics 2025
Computer vision › 3D vision
point cloud processing
0.912025
HiMo: High-Speed Objects Motion Compensation in Point Clouds · IEEE Trans. Robotics 2025
Computer vision › 3D vision › scene flow estimation
self-supervised scene flow
0.912025
HiMo: High-Speed Objects Motion Compensation in Point Clouds · IEEE Trans. Robotics 2025
Geometric modeling and processing
point cloud processing
0.312026
HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint) · AAAI 2026
Computer vision › 3D vision
3d scene understanding
0.312025
SSF: Sparse Long-Range Scene Flow for Autonomous Driving · ICRA 2025
Robotics › Autonomous driving
perception
0.312025
SSF: Sparse Long-Range Scene Flow for Autonomous Driving · ICRA 2025

Methods — techniques the papers use, named apart from their topics

scene flow estimation · 2.9sparse feature fusion · 0.9sparse convolution · 0.9self-supervised learning · 0.9
YearPublicationVenuePosition
2026 HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint)
abstract
LiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for non-ego motion compensation, correcting the representation of dynamic objects in point clouds. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles.
Qingwen Zhang, Ajinkya Khoche, Yi Yang 0095, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt
AAAI2
2026 BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
abstract
Zero-shot 3D object classification is crucial for real-world applications like autonomous driving, however it is often hindered by a significant domain gap between the synthetic data used for training and the sparse, noisy LiDAR scans encountered in the real-world. Current methods trained solely on synthetic data fail to generalize to outdoor scenes, while those trained only on real data lack the semantic diversity to recognize rare or unseen objects. We introduce BlendCLIP, a multimodal pretraining framework that bridges this synthetic-to-real gap by strategically combining the strengths of both domains. We first propose a pipeline to generate a large-scale dataset of object-level triplets—consisting of a point cloud, image, and text description—mined directly from real-world driving data and human annotated 3D boxes. Our core contribution is a curriculum-based data mixing strategy that first grounds the model in the semantically rich synthetic CAD data before progressively adapting it to the specific characteristics of real-world scans. Our experiments show that our approach is highly label-efficient: introducing as few as 1.5% real-world samples per batch into training boosts zero-shot accuracy on the nuScenes benchmark by 27%. Consequently, our final model achieves state-of-the-art performance on challenging outdoor datasets like nuScenes and TruckScenes, improving over the best prior method by 19.3% on nuScenes, while maintaining strong generalization on diverse synthetic benchmarks. Our findings demonstrate that effective domain adaptation, not full-scale real-world annotation, is the key to unlocking robust open-vocabulary 3D perception. Our code and dataset will be released upon acceptance on https://github.com/kesu1/BlendCLIP.
Ajinkya Khoche, Gergo László Nagy, Maciej Wozniak 0001, Thomas Gustafsson, Patric Jensfelt
WACV1
2026 Correcting and Quantifying Systematic Errors in 3D Box Annotations for Autonomous Driving
abstract
Accurate ground truth annotations are critical to supervised learning and evaluating the performance of autonomous vehicle systems. These vehicles are typically equipped with active sensors, such as LiDAR, which scan the environment in predefined patterns. 3D box annotation based on data from such sensors is challenging in dynamic scenarios, where objects are observed at different timestamps, hence different positions. Without proper handling of this phenomenon, systematic errors are prone to being introduced in the box annotations. Our work is the first to discover such annotation errors in widely used, publicly available datasets. Through our novel offline estimation method, we correct the annotations so that they follow physically feasible trajectories and achieve spatial and temporal consistency with the sensor data. For the first time, we define metrics for this problem; and we evaluate our method on the Argoverse 2, MAN TruckScenes, and our proprietary datasets. Our approach increases the quality of box annotations by more than 17% in these datasets. Furthermore, we quantify the annotation errors in them and find that the original annotations are misplaced by up to 2.5 m, with highly dynamic objects being the most affected. Finally, we test the impact of the errors in benchmarking and find that the impact is larger than the improvements that state-of-the-art methods typically achieve w.r.t. the previous state-of-the-art methods; showing that accurate annotations are essential for correct interpretation of performance. Our code is available at https://github.com/alexandre-justo-miro/annotation-correction-3D-boxes.
Alexandre Justo Miro, Ludvig af Klinteberg, Bogdan Timus, Aron Asefaw, Ajinkya Khoche, Thomas Gustafsson, Sina Sharif Mansouri, Masoud Daneshtalab
WACV5
2025 SSF: Sparse Long-Range Scene Flow for Autonomous Driving
abstract
Scene flow enables an understanding of the motion characteristics of the environment in the 3D world. It gains particular significance in the long-range, where object-based perception methods might fail due to sparse observations far away. Although significant advancements have been made in scene flow pipelines to handle large-scale point clouds, a gap remains in scalability with respect to long-range. We attribute this limitation to the common design choice of using dense feature grids, which scale quadratically with range. In this paper, we propose Sparse Scene Flow (SSF), a general pipeline for long-range scene flow, adopting a sparse convolution based backbone for feature extraction. This approach introduces a new challenge: a mismatch in size and ordering of sparse feature maps between time-sequential point scans. To address this, we propose a sparse feature fusion scheme, that augments the feature maps with virtual voxels at missing locations. Additionally, we propose a range-wise metric that implicitly gives greater importance to faraway points. Our method, SSF, achieves state-of-the-art results on the Argoverse2 dataset, demonstrating strong performance in long-range scene flow estimation. Our code is open-sourced at https://github.com/KTH-RPL/SSF.git.
Ajinkya Khoche, Qingwen Zhang, Laura Pereira Sánchez, Aron Asefaw, Sina Sharif Mansouri, Patric Jensfelt
ICRA1
2025 HiMo: High-Speed Objects Motion Compensation in Point Clouds
abstract
LiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for non-ego motion compensation, correcting the representation of dynamic objects in point clouds. During the development of HiMo, we observed that existing self-supervised scene flow estimators often produce degenerate or inconsistent estimates under high-speed distortion. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. Since well-established motion distortion metrics are absent in the literature, we introduce two evaluation metrics: compensation accuracy at a point level and shape similarity of objects. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles. Our findings show that HiMo improves the geometric consistency and visual fidelity of dynamic objects in LiDAR point clouds, benefiting downstream tasks such as semantic segmentation and 3D detection. See https://kin-zhang.github.io/HiMo for more details.
Qingwen Zhang, Ajinkya Khoche, Yi Yang 0095, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt
IEEE Trans. Robotics2
2024 Towards Long-Range 3D Object Detection for Autonomous Vehicles
abstract
3D object detection at long-range is crucial for ensuring the safety and efficiency of self-driving vehicles, allowing them to accurately perceive and react to objects, obstacles, and potential hazards from a distance. But most current state-of-the-art LiDAR based methods are range limited due to sparsity at long-range, which generates a form of domain gap between points closer to and farther away from the ego vehicle. Another related problem is the label imbalance for faraway objects, which inhibits the performance of Deep Neural Networks at long-range. To address the above limitations, we investigate two ways to improve long-range performance of current LiDAR-based 3D detectors. First, we combine two 3D detection networks, referred to as range experts, one specializing at near to mid-range objects, and one at long-range 3D detection. To train a detector at long-range under a scarce label regime, we further weigh the loss according to the labelled point’s distance from ego vehicle. Second, we augment LiDAR scans with virtual points generated using Multimodal Virtual Points (MVP), a readily available image-based depth completion algorithm. Our experiments on the long-range Argoverse2 (AV2) dataset indicate that MVP is more effective in improving long range performance, while maintaining a straightforward implementation. On the other hand, the range experts offer a computationally efficient and simpler alternative, avoiding dependency on image-based segmentation networks and perfect camera-LiDAR calibration.
Ajinkya Khoche, Laura Pereira Sánchez, Nazre Batool, Sina Sharif Mansouri, Patric Jensfelt
IV1