VLDB 2026 Research / reviewers in the wild / expert
Kourosh Khoshelham
dblp:87/7132
· DBLP profile ↗
25ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0001-6639-1727ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 3Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Out-of-distribution detection in 3D applications: A reviewabstractThe ability to detect objects that are not prevalent in the training set is a critical capability in many 3D applications, including autonomous driving. Machine learning methods for object recognition often assume that all object categories encountered during inference belong to a closed set of classes present in the training data. This assumption limits generalization to the real world, as objects not seen during training may be misclassified or entirely ignored. As part of reliable AI, OOD detection identifies inputs that deviate significantly from the training distribution. This paper provides a comprehensive overview of OOD detection within the broader scope of trustworthy and uncertain AI. We begin with key use cases across diverse domains, introduce benchmark datasets spanning multiple modalities, and discuss evaluation metrics. Next, we present a comparative analysis of OOD detection methodologies, exploring model structures, uncertainty indicators, and distributional distance taxonomies, alongside uncertainty calibration techniques. Finally, we highlight promising research directions, including adversarially robust OOD detection and failure identification, particularly relevant to 3D applications. The paper offers both theoretical and practical insights into OOD detection, showcasing emerging research opportunities such as 3D vision integration. These insights help new researchers navigate the field more effectively, contributing to the development of reliable, safe, and robust AI systems. Zizhao Li, Xueyang Kang, Joseph West, Kourosh Khoshelham |
Neurocomputing | 4 |
| 2025 | SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionabstractLidar point cloud synthesis based on generative models offers a promising solution to augment deep learning pipelines, particularly when real-world data is scarce or lacks diversity. By enabling flexible object manipulation, this synthesis approach can significantly enrich training datasets and enhance discriminative models. However, existing methods focus on unconditional lidar point cloud generation, overlooking their potential for real-world applications. In this paper, we propose SG-LDM, a Semantic-Guided Lidar Diffusion Model that employs latent alignment to enable robust semantic-to-lidar synthesis. By directly operating in the native lidar space and leveraging explicit semantic conditioning, SG-LDM achieves state-of-the-art performance in generating high-fidelity lidar point clouds guided by semantic labels. Moreover, we propose the first diffusion-based lidar translation framework based on SG-LDM, which enables cross-domain translation as a domain adaptation strategy to enhance downstream perception performance. Systematic experiments demonstrate that SG-LDM significantly outperforms existing lidar diffusion models and the proposed lidar translation framework further improves data augmentation performance in the downstream lidar segmentation task. Zhengkang Xiang, Zizhao Li, Amir Khodabandeh, Kourosh Khoshelham |
ICCV | 4 |
| 2025 | Multi-view Geometry-Aware Diffusion Transformer for Novel View Synthesis of Indoor ScenesabstractRecent progress in novel view synthesis of indoor scenes using diffusion models has attracted significant attention, particularly for generating desired poses from a source image. Existing methods can produce plausible views near the input view, but they often fail to extrapolate views far beyond the input perspective. Additionally, achieving a multiview consistent diffusion model typically requires training of computationally intensive 3D priors, limiting scalability to long-range generation. In this paper, we present a transformer-based latent diffusion model that leverages view geometry constraints, including explicitly warped feature maps of the input view as the denoised target view and a conditioning combination of epipolar-weighted source image feature map, Plücker raymap, and camera poses. This approach allows for the extrapolation of consistent novel views, both semantically and geometrically, over long-range trajectories in a single-shot manner. Our model is evaluated on two indoor datasets, ScanNet and RealEstate10K, using a diverse set of metrics for view quality and consistency evaluation. Experimental results demonstrate the superiority of our approach over existing models, showcasing its potential for semantically and geometrically consistent novel view synthesis, scalable in video generation applications. Xueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh Khoshelham |
IJCNN | 4 |
| 2025 | Look Beyond: Two-Stage Scene View Generation via Panorama and Video DiffusionabstractNovel view synthesis (NVS) from a single image is highly ill-posed due to large unobserved regions, especially for views that deviate significantly from the input. While existing methods focus on consistency between the source and generated views, they often fail to maintain coherence and correct view alignment across long-range or looped trajectories. We propose a model that addresses this by decomposing single-view NVS into a 360-degree scene extrapolation followed by novel view interpolation. This design ensures long-term view and scene consistency by conditioning on keyframes extracted and warped from a generated panoramic representation. In the first stage, a panorama diffusion model learns the scene prior from the input perspective image. Perspective keyframes are then sampled and warped from the panorama and used as anchor frames in a pre-trained video diffusion model, which generates novel views through a proposed spatial noise diffusion process. Compared to the prior work, our method produces globally consistent novel views-even in loop-closure scenarios, while enabling flexible camera control. Experiments on diverse scene datasets demonstrate that our approach outperforms existing methods in generating coherent views along user-defined trajectories. Our implementation is available at https://github.com/YiGuYT/LookBeyond. Xueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh Khoshelham |
ACM Multimedia | 4 |
| 2025 | Learning geometric invariant features for classification of vector polygons with graph message-passing neural networkabstractAbstract Geometric shape classification of vector polygons remains a challenging task in spatial analysis. Previous studies have primarily focused on deep learning approaches for rasterized vector polygons, while the study of discrete polygon representations and corresponding learning methods remains underexplored. In this study, we investigate a graph-based representation of vector polygons and propose a simple graph message-passing framework, PolyMP, along with its densely self-connected variant, PolyMP-DSC, to learn more expressive and robust latent representations of polygons. This framework hierarchically captures self-looped graph information and learns geometric-invariant features for polygon shape classification. Through extensive experiments, we demonstrate that combining a permutation-invariant graph message-passing neural network with a densely self-connected mechanism achieves robust performance on benchmark datasets, including synthetic glyphs and real-world building footprints, outperforming several baseline methods. Our findings indicate that PolyMP and PolyMP-DSC effectively capture expressive geometric features that remain invariant under common transformations, such as translation, rotation, scaling, and shearing, while also being robust to trivial vertex removals. Furthermore, we highlight the strong generalization ability of the proposed approach, enabling the transfer of learned geometric features from synthetic glyph polygons to real-world building footprints. Zexian Huang, Kourosh Khoshelham, Martin Tomko 0001 |
GeoInformatica | 2 |
| 2025 | Self-supervised monocular depth estimation via joint attention and intelligent mask loss
Shuguo Pan, Kourosh Khoshelham |
Mach. Vis. Appl. | 4 |
| 2024 | Equi-GSPR: Equivariant SE(3) Graph Network Model for Sparse Point Cloud Registration
Xueyang Kang, Zhaoliang Luan, Kourosh Khoshelham, Bing Wang 0013 |
ECCV (4) | 3 |
| 2024 | Synthetic lidar point cloud generation using deep generative models for improved driving scene object recognitionabstractThe imbalanced distribution of different object categories poses a challenge for training accurate object recognition models in driving scenes. Supervised machine learning models trained on imbalanced data are biased and easily overfit the majority classes, such as vehicles and pedestrians, which appear more frequently in driving scenes. We propose a novel data augmentation approach for object recognition in lidar point cloud of driving scenes, which leverages probabilistic generative models to produce synthetic point clouds for the minority classes and complement the original imbalanced dataset. We evaluate five generative models based on different statistical principles, including Gaussian mixture model, variational autoencoder, generative adversarial network, adversarial autoencoder and the diffusion model. Experiments with a real-world autonomous driving dataset show that the synthetic point clouds generated for the minority classes by the Latent Generative Adversarial Network result in significant improvement of object recognition performance for both minority and majority classes. The codes are available at https://github.com/AAAALEX-XIANG/Synthetic-Lidar-Generation. Zhengkang Xiang, Zexian Huang, Kourosh Khoshelham |
Image Vis. Comput. | 3 |
| 2024 | 3-D Building Instance Extraction From High-Resolution Remote Sensing Images and DSM With an End-to-End Deep Neural NetworkabstractThree-dimensional (3D) building models play a vital role in numerous applications including urban planning and smart cities. Recent 3D building modeling methods either rely heavily on available manaually-collected footprint reference or hardly reach real automation on par with manual editing. To approach the automated extraction of instance-level 3D buildings at Level of Detail (LoD) 1, we introduce an innovative end-to-end 3D building instance segmentation model. This model predicts accurate contours and heights of individual buildings simultaneously using ortho-rectified high-resolution remote sensing images and Digital Surface Models (DSMs), getting rid of additional reference data and impirical parameter settings. Firstly, we propose an Anchor-Free Multi-head building extraction network (AFM) tailored for extracting 2D building contours. AFM incorporates a full-resolution, long-range correlation boosted global mask prediction branch along with anchor-free bounding box generation, as well as a newly developed online hard sample mining (OHSM) training procedure based on uncertainty analysis to emphasize error-prone positions in locating building contours. Subsequently, we incorporate a height prediction component to AFM in order to derive accurate building height information, thus creating the comprehensive 3D building extraction model referred to as AFM-3D. The two-stage AFM-3D operates by initially predicting 3D cube proposals, followed by generating refined 3D prismatic models (LoD1 models) for each proposal. Thorough experimentation across different datasets demonstrates the superior performance of AFM and AFM-3D. A significant enhancement of 6.4% quality score is observed on the urban 3D dataset in comparison to recent methods. In addition to the proposed novel methodology, we compare anchor-based and anchor-free bounding box generation mechanisms for remote sensing data, explore pixel-based and contour-based segmentation strategies, evaluate learning-based and empirical height estimation methods, and discuss the indispensability of DSM data in 3D building instance extraction. These analyses yield valuable insights that contribute to the progression of 3D building extraction research. Dawen Yu, Shunping Ji, Shiqing Wei, Kourosh Khoshelham |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Aligning the real and the virtual world: Mixed reality localisation using learning-based 3D-3D model registration
Marko Radanovic, Kourosh Khoshelham, Clive S. Fraser |
Adv. Eng. Informatics | 2 |
| 2023 | MediNet: transfer learning approach with MediNet medical visual database
Hatice Catal Reis, Veysel Turk, Kourosh Khoshelham, Serhat Kaya |
Multim. Tools Appl. | 3 |
| 2022 | Augmented Reality in Education and Remote SensingabstractAugmented Reality (AR) is a technology that has attracted a lot of attention in recent times, especially, after PokemonGo and quite recently with the gaining popularity of metaverse. AR is not a new technology but its development to maturity has always been dependent on many other technologies like computer vision, mobile processing, optics, etc. We are at that stage in the development cycle of AR where all other technologies not only support but can benefit from its use. This research presents one such innovation in the domain of education - for school children or enthusiasts about geospatial technologies such as remote sensing. In this paper, we have discussed the idea, methodology, and demonstration of one such application which establishes a future trend for use of AR in the education of Geospatial Technology. Different types of remote sensing data, their acquisition method, satellite technology, and applications area are explained using AR for a more interactive experience. Learning about geospatial technology will be helpful as the new frontier of Geoinformation is enhancing everyone's day-to-day life and will help in boosting interest in this field. Abhishek Rai, Kamal Jain, Kourosh Khoshelham |
IGARSS | 4 |
| 2022 | A review of augmented reality visualization methods for subsurface utilities
Mohamed Zahlan Abdul Muthalif, Davood Shojaei, Kourosh Khoshelham |
Adv. Eng. Informatics | 3 |
| 2022 | Corrigendum to "A review of augmented reality visualization methods for subsurface utilities" [Adv. Eng. Inf. 51 (2022) 101498]
Mohamed Zahlan Abdul Muthalif, Davood Shojaei, Kourosh Khoshelham |
Adv. Eng. Informatics | 3 |
| 2020 | A Multi-Camera Tracker for Monitoring Pedestrians in Enclosed EnvironmentsabstractMulti-camera pedestrians tracking is a challenging computer vision task. We propose a multi-camera tracker for monitoring pedestrians in an enclosed shopping environment. We assess the performance of the multi-camera tracker in a case study, tracking customers in a food and speciality market hall. Our multi-camera tracker tracks customers' walking between the stalls in the market. The information is useful for market management, visitor safety, and other potential application areas. Xusheng Wu, Stephan Winter 0001, Kourosh Khoshelham |
ICTAI | 3 |
| 2020 | Landmark Graph-Based Indoor LocalizationabstractIndoor localization is important for a variety of applications, such as location-based services, mobile social networks, and emergency response. Fusing spatial information is an effective way to achieve accurate indoor localization with little or with no need for extra hardware. However, the existing indoor localization methods that make use of spatial information are either computationally expensive or sensitive to the completeness of landmarks. In this article, we propose a novel, low-cost, high-accuracy indoor localization method based on a landmark graph. The experimental results show that the proposed method outperforms the state-of-the-art methods. Fuqiang Gu, Shahrokh Valaee, Kourosh Khoshelham, Jianga Shang, Rui Zhang 0003 |
IEEE Internet Things J. | 3 |
| 2020 | MS-RRFSegNet: Multiscale Regional Relation Feature Segmentation Network for Semantic Segmentation of Urban Scene Point CloudsabstractSemantic segmentation is one of the fundamental tasks in understanding and applying urban scene point clouds. Recently, deep learning has been introduced to the field of point cloud processing. However, compared to images that are characterized by their regular data structure, a point cloud is a set of unordered points, which makes semantic segmentation a challenge. Consequently, the existing deep learning methods for semantic segmentation of point cloud achieve less success than those applied to images. In this article, we propose a novel method for urban scene point cloud semantic segmentation using deep learning. First, we use homogeneous supervoxels to reorganize raw point clouds to effectively reduce the computational complexity and improve the nonuniform distribution. Then, we use supervoxels as basic processing units, which can further expand receptive fields to obtain more descriptive contexts. Next, a sparse autoencoder (SAE) is presented for feature embedding representations of the supervoxels. Subsequently, we propose a regional relation feature reasoning module (RRFRM) inspired by relation reasoning network and design a multiscale regional relation feature segmentation network (MS-RRFSegNet) based on the RRFRM to semantically label supervoxels. Finally, the supervoxel-level inferences are transformed into point-level fine-grained predictions. The proposed framework is evaluated in two open benchmarks (Paris-Lille-3D and Semantic3D). The evaluation results show that the proposed method achieves competitive overall performance and outperforms other related approaches in several object categories. An implementation of our method is available at: https://github.com/HiphonL/MS_RRFSegNet. Chongcheng Chen, Lina Fang, Kourosh Khoshelham, Guixi Shen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | ZeeFi: Zero-Effort Floor Identification with Deep Learning for Indoor LocalizationabstractThe knowledge of the floor-level location of a user in a multi-storey building is important for many applications, especially for emergency response. Existing floor identification systems suffer from a variety of limitations such as low accuracy, the need for a time-consuming site survey, assumption of user encounters, knowledge of the initial floor, and/or poor applicability. In this paper, we propose a novel, zero-effort, deep learning-based floor identification system, called ZeeFi. The proposed system uses the widely-available smartphone sensing to identify on which floor a user is located. By recognizing the ground floor automatically, the proposed system does not require site survey, initial floor knowledge, and other assumptions. To achieve accurate floor identification performance, we have developed a deep learning-based method. Experimental results show that the proposed system outperforms the state-of-the-art systems, and is very promising for large-scale deployment. Fuqiang Gu, Jörg Blankenbach, Kourosh Khoshelham, Jan Grottke, Shahrokh Valaee |
GLOBECOM | 3 |
| 2018 | Locomotion Activity Recognition Using Stacked Denoising AutoencodersabstractLocomotion activity recognition (LAR) is important for a number of applications, such as indoor localization, fitness tracking, and aged care. Existing methods usually use handcrafted features, which requires expert knowledge and is laborious, and the achieved result might still be suboptimal. To relieve the burden of designing and selecting features, we propose a deep learning method for LAR by using data from multiple sensors available on most smart devices. Experimental results show that the proposed method, which learns useful features automatically, outperforms conventional classifiers that require the hand-engineering of features. We also show that the combination of sensor data from four sensors (accelerometer, gyroscope, magnetometer, and barometer) achieves a higher accuracy than other combinations or individual sensors. Fuqiang Gu, Kourosh Khoshelham, Shahrokh Valaee, Jianga Shang, Rui Zhang 0003 |
IEEE Internet Things J. | 2 |
| 2017 | Omnidirectional visual-inertial odometry using multi-state constraint Kalman filterabstractWe present an Omnidirectional Visual-Inertial Odometry (OVIO) approach based on Multi-State Constraint Kalman Filtering (MSCKF) to estimate the ego-motion of a moving platform. Instead of considering visual measurements on image plane, we use individual planes for each point that are tangent to the unit sphere and normal to the corresponding measurement ray. This way, we combine spherical images captured by omnidirectional camera with inertial measurements within the filtering method MSCKF. The key hypothesis of OVIO is that a wider field of view allows incorporating more visual features from the surrounding environment, thereby improving the accuracy and robustness of the motion estimation. Moreover, by using an omnidirectional camera, it is less likely to end up in a situation where there is not enough texture. We provide an evaluation of OVIO using synthetic and real video sequences captured by a fish-eye camera, and compare the performance with MSCKF using a perspective camera. The results show the superior performance of the proposed OVIO. Milad Ramezani, Kourosh Khoshelham, Laurent Kneip |
IROS | 2 |
| 2017 | Locomotion activity recognition: A deep learning approachabstractHuman activity recognition is important for a large number of applications including indoor localization. Existing methods usually involve manually-designed features, which require expert knowledge and are laborious. Also, previous works use only the accelerometer for activity recognition, which may fail to recognize some complex activities. In this paper, we propose a deep learning-based method for locomotion activity recognition by using the combination of data from multiple smartphone built-in sensors. Eight types of locomotion activities are identified including the new `False Motion' activity introduced for the first time in this work. Experimental results show that the proposed method, which learns useful features automatically, outperforms conventional classifiers that require hand-engineering of features. Also, using data from multiple sensors helps to improve recognition accuracy by about 10% compared to that using accelerometer data only. Fuqiang Gu, Kourosh Khoshelham, Shahrokh Valaee |
PIMRC | 2 |
| 2016 | Mapping Indoor Spaces by Adaptive Coarse-to-Fine Registration of RGB-D DataabstractIn this letter, we present an adaptive coarse-to-fine registration method for 3-D indoor mapping using RGB-D data. We weight the 3-D points based on the theoretical random error of depth measurements and introduce a novel disparity-based model for an accurate and robust coarse-to-fine registration. Some feature extraction methods required by the method are also presented. First, our method exploits both visual and depth information to compute the initial transformation parameters. We employ scale-invariant feature transformation for extracting, detecting, and matching 2-D visual features, and their associated depth values are used to perform coarse registration. Then, we use an image-based segmentation technique for detecting regions in the RGB images. Their associated 3-D centroid and the correspondent disparity values are used to refine the initial transformation parameters. Finally, the loop-closure detection and a global adjustment of the complete sequence data are used to recognize when the camera has returned to a previously visited location and minimize the registration errors. The effectiveness of the proposed method is demonstrated with the Kinect data set. The experimental results show that the proposed method can properly map the indoor environment with a relative and absolute accuracy value of around 3-5 cm, respectively. Daniel Rodrigues dos Santos, Marcos A. Basso, Kourosh Khoshelham, Elizeu de Oliveira, Nadisson L. Pavan, George Vosselman |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Localized Registration of Point Clouds of Botanic TreesabstractA global registration is often insufficient for estimating dendrometric characteristics of trees because individual branches of the same tree may exhibit different positions between two scanning procedures. Therefore, we introduce a localized approach to register point clouds of botanic trees. Given two roughly registered point clouds PC1and PC2of a tree, we apply a skeletonization method to both point clouds. Based on these two skeletons, initial correspondences between branch segments of both point clouds are established to estimate local transformation parameters. The transformation estimation relies on minimizing the distance between the points in PC1and the skeleton of PC2. The performance of the method is demonstrated on two example trees. It is shown that significant improvements can be achieved for the registration of fine branches. These improvements are quantified as the residual point-to-line distances before and after the localized fine registration. In our experiment, the residual error after the local registration is on an average of 5 mm over 90 skeleton segments, which is about three times smaller than the average residual error of the initial rough registration. Alexander Bucksch, Kourosh Khoshelham |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2013 | Segment-Based Classification of Damaged Building Roofs in Aerial Laser Scanning DataabstractIdentifying damaged buildings after natural disasters such as earthquake is important for the planning of recovery actions. We present a segment-based approach to classifying damaged building roofs in aerial laser scanning data. A challenge in the supervised classification of point segments is the generation of training samples, which is difficult because of the complexity of interpreting point clouds. We evaluate the performance of three different classifiers trained with a small set of training samples and show that feature selection improves the training and the accuracy of the resulting classification. When trained with 50 training samples, a linear discriminant classifier using a subset of six features reaches a classification accuracy of 85%. Kourosh Khoshelham, Sander Oude Elberink, Sudan Xu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2010 | Automated localization of a laser scanner in indoor environments using planar objectsabstractA method is presented for automated localization of a laser scanner in indoor environments by matching planar features extracted from range data. Plane correspondences are used in a linear least-squares adjustment model to estimate the relative scanner positions in consecutive scans. The performance of the method is demonstrated using two datasets of building interiors. Accuracy assessment of the computed positions shows localization errors of a few centimeters. Kourosh Khoshelham |
IPIN | 1 |