VLDB 2026 Research / reviewers in the wild / expert
Waqas Sultani
dblp:18/8614
· DBLP profile ↗
26ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0002-9322-0728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-View Meets Diffusion: Aerial Image Synthesis with Geometry and Text GuidanceabstractAerial imagery analysis is critical for many research fields. However, obtaining frequent high-quality aerial images is not always accessible due to its high effort and cost requirements. One solution is to use the Ground-to-Aerial (G2A) technique to synthesize aerial images from easily collectible ground images. However, G2A is rarely studied, because of its challenges, including but not limited to, the drastic view changes, occlusion, and range of visibil-ity. In this paper, we present a novel Geometric Preserving Ground-to-Aerial (G2A) image synthesis (GPG2A) model that can generate realistic aerial images from ground images. GPG2A consists of two stages. The first stage predicts the Bird's Eye View (BEV) segmentation (referred to as the BEV layout map) from the ground image. The second stage synthesizes the aerial image from the predicted BEV layout map and text descriptions of the ground image. To train our model, we present a new multimodal cross-view dataset, namely VIGORv2, built upon VIGOR [64] with newly collected aerial images, maps, and text descriptions. Our extensive experiments illustrate that GPG2A synthesizes better geometry-preserved aerial images than existing models. We also present two applications, data augmentation for cross-view geo-localization and sketch-based region search, to further verify the effectiveness of our GPG2A. The code and dataset are available at https://github.com/AhmadArrabi/GPG2A. Ahmad Arrabi, Xiaohan Zhang 0003, Waqas Sultani, Chen Chen 0001, Safwan Wshah |
WACV | 3 |
| 2025 | R2S100K: Road-Region Segmentation Dataset for Semi-supervised Autonomous Driving in the WildabstractAbstract Semantic understanding of roadways is a key enabling factor for safe autonomous driving. However, existing autonomous driving datasets provide well-structured urban roads while ignoring unstructured roadways containing distress, potholes, water puddles, and various kinds of road patches i.e., earthen, gravel etc. To this end, we introduce Road Region Segmentation dataset (R2S100K)—a large-scale dataset and benchmark for training and evaluation of road segmentation in aforementioned challenging unstructured roadways. R2S100K comprises 100K images extracted from a large and diverse set of video sequences covering more than 1000 km of roadways. Out of these 100K privacy respecting images, 14,000 images have fine pixel-labeling of road regions, with 86,000 unlabeled images that can be leveraged through semi-supervised learning methods. Alongside, we present an Efficient Data Sampling based self-training framework to improve learning by leveraging unlabeled data. Our experimental results demonstrate that the proposed method significantly improves learning methods in generalizability and reduces the labeling cost for semantic segmentation tasks. Our benchmark will be publicly available to facilitate future research at https://r2s100k.github.io/ . Muhammad Atif Butt, Hassan Ali 0001, Adnan Qayyum, Waqas Sultani, Ala I. Al-Fuqaha, Junaid Qadir 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Leveraging sparse annotations for leukemia diagnosis on the large leukemia dataset
Abdul Rehman 0012, Talha Meraj, Aiman Mahmood Minhas, Ayisha Imran, Mohsen Ali, Waqas Sultani, Mubarak Shah |
Medical Image Anal. | 6 |
| 2024 | Few-Shot Domain Adaptive Object Detection for Microscopic Images
Sumayya Inayat, Nimra Dilawar, Waqas Sultani, Mohsen Ali |
MICCAI (12) | 3 |
| 2024 | A Large-Scale Multi Domain Leukemia Dataset for the White Blood Cells Detection with Morphological Attributes for Explainability
Abdul Rehman 0012, Talha Meraj, Aiman Mahmood Minhas, Ayisha Imran, Mohsen Ali, Waqas Sultani |
MICCAI (3) | 6 |
| 2024 | Image and Object Geo-Localization
Daniel Wilson 0001, Xiaohan Zhang 0003, Waqas Sultani, Safwan Wshah |
Int. J. Comput. Vis. | 3 |
| 2024 | GeoDTR+: Toward Generic Cross-View Geolocalization via Geometric DisentanglementabstractCross-View Geo-Localization (CVGL) estimates the location of a ground image by matching it to a geo-tagged aerial image in a database. Recent works achieve outstanding progress on CVGL benchmarks. However, existing methods still suffer from poor performance in cross-area evaluation, in which the training and testing data are captured from completely distinct areas. We attribute this deficiency to the lack of ability to extract the geometric layout of visual features and models' overfitting to low-level details. Our preliminary work (Zhang et al. 2022) introduced a Geometric Layout Extractor (GLE) to capture the geometric layout from input features. However, the previous GLE does not fully exploit information in the input feature. In this work, we propose GeoDTR+ with an enhanced GLE module that better models the correlations among visual features. To fully explore the LS techniques from our preliminary work, we further propose Contrastive Hard Samples Generation (CHSG) to facilitate model training. Extensive experiments show that GeoDTR+ achieves state-of-the-art (SOTA) results in cross-area evaluation on CVUSA (Workman et al. 2015), CVACT (Liu and Li, 2019), and VIGOR (Zhu et al. 2021) by a large margin (16.44%, 22.71%, and 13.66% without polar transformation) while keeping the same-area performance comparable to existing SOTA. Moreover, we provide detailed analyses of GeoDTR+. Xiaohan Zhang 0003, Waqas Sultani, Chen Chen 0001, Safwan Wshah |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Cross-View Geo-Localization via Learning Disentangled Geometric Layout CorrespondenceabstractCross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and different capturing time between two views. Despite these difficulties, recent works achieve outstanding progress on cross-view geo-localization benchmarks. However, existing methods still suffer from poor performance on the cross-area benchmarks, in which the training and testing data are captured from two different regions. We attribute this deficiency to the lack of ability to extract the spatial configuration of visual feature layouts and models' overfitting on low-level details from the training set. In this paper, we propose GeoDTR which explicitly disentangles geometric information from raw features and learns the spatial correlations among visual features from aerial and ground pairs with a novel geometric layout extractor module. This module generates a set of geometric layout descriptors, modulating the raw features and producing high-quality latent representations. In addition, we elaborate on two categories of data augmentations, (i) Layout simulation, which varies the spatial configuration while keeping the low-level details intact. (ii) Semantic augmentation, which alters the low-level details and encourages the model to capture spatial configurations. These augmentations help to improve the performance of the cross-view geo-localization models, especially on the cross-area benchmarks. Moreover, we propose a counterfactual-based learning process to benefit the geometric layout extractor in exploring spatial information. Extensive experiments show that GeoDTR not only achieves state-of-the-art results but also significantly boosts the performance on same-area and cross-area benchmarks. Our code can be found at https://gitlab.com/vail-uvm/geodtr. Xiaohan Zhang 0003, Waqas Sultani, Safwan Wshah |
AAAI | 3 |
| 2023 | TransVisDrone: Spatio-Temporal Transformer for Vision-based Drone-to-Drone Detection in Aerial VideosabstractDrone-to-drone detection using visual feed has crucial applications, such as detecting drone collisions, detecting drone attacks, or coordinating flight with other drones. However, existing methods are computationally costly, follow non-end-to-end optimization, and have complex multi-stage pipelines, making them less suitable for real-time deployment on edge devices. In this work, we propose a simple yet effective framework, TransVisDrone, that provides an end-to-end solution with higher computational efficiency. We utilize CSPDarkNet-53 network to learn object-related spatial features and VideoSwin model to improve drone detection in challenging scenarios by learning spatio-temporal dependencies of drone motion. Our method achieves state-of-the-art performance on three challenging real-world datasets (Average [email protected]): NPS 0.95, FLDrones 0.75, and AOT 0.80, and a higher throughput than previous methods. We also demonstrate its deployment capability on edge devices and its usefulness in detecting drone-collision (encounter). Project: https://tusharsangam.github.io/TransVisDrone-project-page/ Tushar Sangam, Ishan Rajendrakumar Dave, Waqas Sultani, Mubarak Shah |
ICRA | 3 |
| 2023 | Cross-View Image Sequence Geo-localizationabstractCross-view geo-localization aims to estimate the GPS location of a query ground-view image by matching it to images from a reference database of geo-tagged aerial images. To address this challenging problem, recent approaches use panoramic ground-view images to increase the range of visibility. Although appealing, panoramic images are not readily available compared to the videos of limited Field-Of-View (FOV) images. In this paper, we present the first cross-view geo-localization method that works on a sequence of limited FOV images. Our model is trained end-to-end to capture the temporal structure that lies within the frames using the attention-based temporal feature aggregation module. To robustly tackle different sequences length and GPS noises during inference, we propose to use a sequential dropout scheme to simulate variant length sequences. To evaluate the proposed approach in realistic settings, we present a new large-scale dataset containing ground-view sequences along with the corresponding aerial-view images. Extensive experiments and comparisons demonstrate the superiority of the proposed approach compared to several competitive baselines. Xiaohan Zhang 0003, Waqas Sultani, Safwan Wshah |
WACV | 2 |
| 2023 | Cross-region building counting in satellite imagery using counting consistency
Muaaz Zakria, Hamza Rawal, Waqas Sultani, Mohsen Ali |
Neural Comput. Appl. | 3 |
| 2023 | Fine-Grained Road Quality Monitoring Using Deep LearningabstractExisting solutions for road surface monitoring have assumed that events (both normal and abnormal) have a defined duration, and these methods fail to provide a universal framework that can be used to assess the quality of roads in real-world scenarios where events do not have to be of fixed duration. This article aims to improve road quality assessment systems by overcoming the constraint of fixed window size and taking into account the real-world scenario of variable-length events. First, we annotate a big heterogeneous data set without partitioning it into fixed-size windows. Second, we suggest two distinct approaches for detecting and characterizing anomalies utilizing deep learning architectures comprising Bi-Directional LSTM units. The first strategy is sequence classification (using a many to one correspondence to classify the entire sequence), and the second approach is endpoint detection (classify each time step as a normal or anomalous event using a many to many approach). The solutions presented in this article are intended for use with non-anomalous (normal) signals as well as with four distinct types of anomalies: cat-eyes, manholes, potholes, and speed bumps. Our sequence classification model (Bi-LSTM model) is capable of detecting anomalies with a 97.3% True Positive Rate when the anomalies are considered a positive class. On the other hand, our end-point detection framework is able to mark the exact end-point of anomalous signals with a true positive rate of 90.2% as shown in II. Our dataset and annotation are publicly available at:https://drive.google.com/drive/folders/1Qf-4D6P9Oeu-yw3yc3UY55V7wDCWiumI Ifrah Siddiqui, Suleman Mazhar, Naufil Hassan, Waqas Sultani |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Towards Low-Cost and Efficient Malaria DetectionabstractMalaria, a fatal but curable disease claims hundreds of thousands of lives every year. Early and correct diagnosis is vital to avoid health complexities, however, it depends upon the availability of costly microscopes and trained experts to analyze blood-smear slides. Deep learning-based methods have the potential to not only decrease the burden of experts but also improve diagnostic accuracy on low-cost microscopes. However, this is hampered by the absence of a reasonable size dataset. One of the most challenging aspects is the reluctance of the experts to annotate the dataset at low magnification on low-cost microscopes. We present a dataset to further the research on malaria microscopy over low-cost microscopes at low magnification. Our large-scale dataset consists of images of blood-smear slides from several malaria-infected patients, collected through micro-scopes at two different cost spectrums and multiple magnifications. Malarial cells are annotated for the localization and life-stage classification task on the images collected through the high-cost microscope at high magnification. We design a mechanism to transfer these annotations from the high-cost microscope at high magnification to the low-cost microscope, at multiple magnifications. Multiple object detectors and domain adaptation methods are presented as the baselines. Furthermore, a partially supervised domain adaptation method is introduced to adapt the object-detector to work on the images collected from the low-cost microscope. The dataset is available here: http://im.itu.edu.pk/m5-malaria-dataset/ Waqas Sultani, Wajahat Nawaz, Syed Javed, Muhammad Sohail Danish, Asma Saadia, Mohsen Ali |
CVPR | 1 |
| 2022 | Mapping Temporary Slums From Satellite Imagery Using a Semi-Supervised ApproachabstractOne billion people worldwide are estimated to be living in slums, and documenting and analyzing these regions is a challenging task. As compared to regular slums; the small, scattered and temporary nature of temporary slums makes data collection and labeling tedious and time-consuming. To tackle this challenging problem of temporary slums detection, we present a semi-supervised deep learning segmentation-based approach; with the strategy to detect initial seed images in the zero-labeled data settings. A small set of seed samples (32 in our case) are automatically discovered by analyzing the temporal changes, which are manually labeled to train a segmentation and representation learning module. The segmentation module gathers high dimensional image representations, and the representation learning module transforms image representations into embedding vectors. After that, a scoring module uses the embedding vectors to sample images from a large pool of unlabeled images and generates pseudo-labels for the sampled images. These sampled images with their pseudo-labels are added to the training set to update the segmentation and representation learning modules iteratively. To analyze the effectiveness of our technique, we construct a large geographically marked dataset of temporary slums. This dataset constitutes more than 200 potential temporary slum locations (2.28 square kilometers) found by sieving sixty-eight thousand images from 12 metropolitan cities of Pakistan covering 8000 square kilometers. Furthermore, our proposed method outperforms several competitive semi-supervised semantic segmentation baselines on a similar setting. The code and the dataset will be made publicly available. M. Fasi ur Rehman, Izza Aftab, Waqas Sultani, Mohsen Ali |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A dataset and benchmark for malaria life-cycle classification in thin blood smear images
Qazi Ammar Arshad, Mohsen Ali, Saeed-Ul Hassan, Chen Chen 0001, Ayisha Imran, Ghulam Rasul, Waqas Sultani |
Neural Comput. Appl. | 7 |
| 2022 | Fake visual content detection using two-stream convolutional neural networks
Bilal Yousaf, Waqas Sultani, Arif Mahmood, Junaid Qadir 0001 |
Neural Comput. Appl. | 3 |
| 2021 | Dogfight: Detecting Drones From Drones Videos
Muhammad Waseem Ashraf, Waqas Sultani, Mubarak Shah |
CVPR | 2 |
| 2021 | Human action recognition in drone videos using a few aerial training examples
Waqas Sultani, Mubarak Shah |
Comput. Vis. Image Underst. | 1 |
| 2020 | Single-Shot Retinal Image Enhancement Using Deep Image Priors
Adnan Qayyum, Waqas Sultani, Fahad Shamshad, Junaid Qadir 0001, Rashid Tufail |
MICCAI (5) | 2 |
| 2018 | Real-World Anomaly Detection in Surveillance VideosabstractSurveillance videos are able to capture a variety of realistic anomalies. In this paper, we propose to learn anomalies by exploiting both normal and anomalous videos. To avoid annotating the anomalous segments or clips in training videos, which is very time consuming, we propose to learn anomaly through the deep multiple instance ranking framework by leveraging weakly labeled training videos, i.e. the training labels (anomalous or normal) are at video-level instead of clip-level. In our approach, we consider normal and anomalous videos as bags and video segments as instances in multiple instance learning (MIL), and automatically learn a deep anomaly ranking model that predicts high anomaly scores for anomalous video segments. Furthermore, we introduce sparsity and temporal smoothness constraints in the ranking loss function to better localize anomaly during training. We also introduce a new large-scale first of its kind dataset of 128 hours of videos. It consists of 1900 long and untrimmed real-world surveillance videos, with 13 realistic anomalies such as fighting, road accident, burglary, robbery, etc. as well as normal activities. This dataset can be used for two tasks. First, general anomaly detection considering all anomalies in one group and all normal activities in another group. Second, for recognizing each of 13 anomalous activities. Our experimental results show that our MIL method for anomaly detection achieves significant improvement on anomaly detection performance as compared to the state-of-the-art approaches. We provide the results of several recent deep learning baselines on anomalous activity recognition. The low recognition performance of these baselines reveals that our dataset is very challenging and opens more opportunities for future work. The dataset is available at: http://crcv.ucf.edu/projects/real-world. Waqas Sultani, Chen Chen 0001, Mubarak Shah |
CVPR | 1 |
| 2018 | Automatic Pavement Object Detection Using Superpixel Segmentation Combined With Conditional Random FieldabstractPavement images contain various objects, such as lane-marker, manhole covers, patches, potholes, and curbing. Accurate and robust computer vision algorithms are necessary to detect these various objects that have random shapes, colors, and sizes. In this paper, we have addressed the problem of automatic object detection in pavement images using a unified framework. To detect an object of arbitrary shape in an efficient way, we first divide the image into small consistent regions called superpixels. These superpixels are fast to calculate and preserve object boundaries. We then compute several texture and intensity features within each superpixel. After that, we train support vector machine (SVM) classifier for every feature separately in one-verses-all paradigm. In testing, we first estimate the probability of each superpixel being the part of some object of interest using these SVM classifiers. Since these superpixels' probabilistic scores are independently computed, they do not preserve neighborhood consistency. Therefore, to enforce superpixel neighborhood label consistency, we use contextual optimization technique i.e., conditional random field (CRF). The output of CRF is a pixel-wise binary label map for the objects and background. In addition, due to the lack of any publically available dataset for pavement objects' detection evaluation, we have introduced a new challenging object detection dataset for pavement images. We have performed extensive experiments on this dataset and have obtained encouraging results. Waqas Sultani, Soroush Mokhtari, Hae-Bum Yun |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Automatic action annotation in weakly labeled videos
Waqas Sultani, Mubarak Shah |
Comput. Vis. Image Underst. | 1 |
| 2017 | Unsupervised action proposal ranking through proposal recombination
Waqas Sultani, Mubarak Shah |
Comput. Vis. Image Underst. | 1 |
| 2016 | What If We Do Not have Multiple Videos of the Same Action? - Video Action Localization Using Web ImagesabstractThis paper tackles the problem of spatio-temporal action localization in a video, without assuming the availability of multiple videos or any prior annotations. Action is localized by employing images downloaded from internet using action name. Given web images, we first dampen image noise using random walk and evade distracting backgrounds within images using image action proposals. Then, given a video, we generate multiple spatio-temporal action proposals. We suppress camera and background generated proposals by exploiting optical flow gradients within proposals. To obtain the most action representative proposals, we propose to reconstruct action proposals in the video by leveraging the action proposals in images. Moreover, we preserve the temporal smoothness of the video and reconstruct all proposal bounding boxes jointly using the constraints that push the coefficients for each bounding box toward a common consensus, thus enforcing the coefficient similarity across multiple frames. We solve this optimization problem using variant of two-metric projection algorithm. Finally, the video proposal that has the lowest reconstruction cost and is motion salient is used to localize the action. Our method is not only applicable to the trimmed videos, but it can also be used for action localization in untrimmed videos, which is a very challenging problem. We present extensive experiments on trimmed as well as untrimmed datasets to validate the effectiveness of the proposed approach. Waqas Sultani, Mubarak Shah |
CVPR | 1 |
| 2014 | Human Action Recognition across Datasets by Foreground-Weighted Histogram DecompositionabstractThis paper attempts to address the problem of recognizing human actions while training and testing on distinct datasets, when test videos are neither labeled nor available during training. In this scenario, learning of a joint vocabulary, or domain transfer techniques are not applicable. We first explore reasons for poor classifier performance when tested on novel datasets, and quantify the effect of scene backgrounds on action representations and recognition. Using only the background features and partitioning of gist feature space, we show that the background scenes in recent datasets are quite discriminative and can be used classify an action with reasonable accuracy. We then propose a new process to obtain a measure of confidence in each pixel of the video being a foreground region, using motion, appearance, and saliency together in a 3D MRF based framework. We also propose multiple ways to exploit the foreground confidence: to improve bag-of-words vocabulary, histogram representation of a video, and a novel histogram decomposition based representation and kernel. We used these foreground confidences to recognize actions trained on one data set and test on a different data set. We have performed extensive experiments on several datasets that improve cross dataset recognition accuracy as compared to baseline methods. Waqas Sultani, Imran Saleemi |
CVPR | 1 |
| 2010 | Abnormal Traffic Detection Using Intelligent Driver ModelabstractWe present a novel approach for detecting and localizing abnormal traffic using intelligent driver model. Specifically, we advect particles over video sequence. By treating each particle as a car, we compute driver behavior using intelligent driver model. The behaviors are learned using latent dirichlet allocation and frames are classified as abnormal using likelihood threshold criteria. In order to localize the abnormality; we compute spatial gradients of behaviors and construct Finite Time Lyaponov Field. Finally the region of abnormality is segmented using watershed algorithm. The effectiveness of proposed approach is validated using videos from stock footage websites. Waqas Sultani |
ICPR | 1 |