EDBT 2026 Demo / reviewers in the wild / expert
Syed Zulqarnain Gilani
dblp:43/7407
· DBLP profile ↗
24ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-7448-2327ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RampWatch: An In-the-Wild Dataset and Text-Guided Detection Framework for Recreational VesselsabstractDetecting small, recreational vessels in coastal environments remains a challenge due to complex backgrounds, dynamic lighting conditions, and the scarcity of annotated data for non-commercial maritime traffic. Despite their socio-economic significance, recreational boats are underrepresented in existing datasets and are poorly detected by standard object detectors, particularly in open-vocabulary scenarios. To address this gap, we present RampWatch, an in-the-wild dataset curated from surveillance footage at multiple boat ramps. RampWatch provides instance annotations across 7 categories of recreational vessels, captured under diverse weather, lighting, and occlusion conditions. To benchmark detection in this domain, we introduce YOLO-TG, a novel detection framework that augments YOLOv11 with a text encoder for open-vocabulary recognition and a self-attention module for enhanced spatial reasoning. YOLO-TG adopts a dual-stream design: visual features are extracted via a hierarchical YOLO backbone, while semantic embeddings from natural language prompts are encoded by a frozen language encoder. These modalities are fused via lightweight cross-modal attention, enabling text-guided detection without retraining. YOLO-TG achieves a +12% relative improvement in mAP@50–95 over strong YOLOv11 baselines on RampWatch, and demonstrates robust cross-domain generalization, with gains of +22% on the Singapore Maritime Dataset and +4.3% on the Split Port Ship Classification Dataset. These results highlight the effectiveness of cross-modal grounding and domain-specific datasets for advancing open-world maritime surveillance. Malik Muhammad Asim, Claire B. Smallwood, Abdullah Tariq, Johnny Lo, Syed Zulqarnain Gilani |
WACV | 5 |
| 2025 | BiFuseNet: A Multimodal Network for Estimating Blood Alcohol Concentration via Bidirectional Hierarchical FusionabstractDrunk driving remains a significant public safety challenge, demanding innovative alternatives to conventional methods such as field sobriety tests and breathalysers. Estimating a driver's level of intoxication through facial cues is particularly challenging due to the subtle and person-specific nature of alcohol-induced behaviours. In this paper, we present BiFuseNet, a 3D spatio-temporal multi-modal network designed to classify alcohol impairment levels into three categories: sober, moderate, and severe. Unlike prior approaches that rely on either uni-modal RGB video or hand-crafted facial features, our method exploits complementary physiological cues from RGB and infrared (IR) facial videos. We introduce a Bi-directional Hierarchical Fusion (BiHF) module that applies cross-attention mechanisms at multiple semantic levels of our BiFuseNet, including early, middle, and late feature stages. This enables deep integration of modality-specific signals across varying temporal and spatial contexts. To capture both short-term facial movements and sustained facial dynamics, we implement a sliding window strategy that samples over 30 frames across ten-minute recordings. Extensive experiments on a public dataset demonstrate that BiFuseNet outperforms uni-modal and traditional fusion baselines, achieving a classification accuracy of 88.41% and an AUC-ROC of 0.91, establishing a new state of the art in estimating blood alcohol concentration. Abdullah Tariq, Arooba Maqsood, Martin Masek, Syed Zulqarnain Gilani |
ICMI | 4 |
| 2025 | From Pixels to Prognosis: A Multi-modal Attention-Based Framework for Visceral Adipose Tissue Estimation
Arooba Maqsood, Afsah Saleem, Marc Sim, David Suter, Simone Radavelli-Bagatini, Jonathan M. Hodgson, Richard Prince, Kun Zhu 0028, William D. Leslie, John T. Schousboe, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (15) | 12 |
| 2025 | A Hybrid Contrastive Ordinal Regression Method for Advancing Disease Severity Assessment in Imbalanced Medical Datasets
Afsah Saleem, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (13) | 3 |
| 2024 | A Hybrid CNN-Transformer Feature Pyramid Network for Granular Abdominal Aortic Calcification Detection from DXA Images
Zaid Ilyas, Afsah Saleem, David Suter, John T. Schousboe, William D. Leslie, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (11) | 7 |
| 2024 | Estimating Blood Alcohol Level Through Facial Features for Driver Impairment AssessmentabstractDrunk driving-related road accidents contribute significantly to the global burden of road injuries. Addressing alcohol-related harm, particularly during safety-critical activities like driving, requires real-time monitoring of an individual’s blood alcohol concentration (BAC). We devise an in-vehicle machine learning system that harnesses standard commercial RGB cameras to predict critical levels of BAC. Our system can detect instances of alcohol intoxication impairment as subtle as 0.05 g/dL (WHO recommended legal limit for driving), with an accuracy of 75%, by leveraging the physiological manifestations of alcohol intoxication on a driver’s face. This system holds great promise for improving road safety. In tandem, we have compiled a data set of 60 subjects engaged in simulated driving scenarios, spanning three levels of alcohol intoxication. These scenarios were captured and divided into video segments labeled "sober","low’, and "severe" Alcohol Intoxication Impairment (AII), constituting the basis for evaluating our system’s performance. To the best of our knowledge, this study is the first to create a large-scale real-life dataset of alcohol intoxication and assess intoxication levels using an off-the-shelf RGB camera to detect drunk driving. Ensiyeh Keshtkaran, Brodie von Berg, Grant Regan, David Suter, Syed Zulqarnain Gilani |
WACV | 5 |
| 2023 | SCOL: Supervised Contrastive Ordinal Loss for Abdominal Aortic Calcification Scoring on Vertebral Fracture Assessment Scans
Afsah Saleem, Zaid Ilyas, David Suter, Ghulam M. Hassan, Siobhan Reid, John T. Schousboe, Richard Prince, William D. Leslie, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (6) | 10 |
| 2023 | Unsupervised Learning for Maximum Consensus Robust Fitting: A Reinforcement Learning ApproachabstractRobust model fitting is a core algorithm in several computer vision applications. Despite being studied for decades, solving this problem efficiently for datasets that are heavily contaminated by outliers is still challenging: due to the underlying computational complexity. A recent focus has been on learning-based algorithms. However, most of these approaches are supervised (which require a large amount of labelled training data). In this paper, we introduce a novel unsupervised learning framework: that learns to directly (without labelled data) solve robust model fitting. Moreover, unlike other learning-based methods, our work is agnostic to the underlying input features, and can be easily generalized to a wide variety of LP-type problems with quasi-convex residuals. We empirically show that our method outperforms existing (un)supervised learning approaches, and also achieves competitive results compared to traditional (non-learning-based) methods. Our approach is designed to try to maximise consensus (MaxCon), similar to the popular RANSAC. The basis of our approach, is to adopt a Reinforcement Learning framework. This requires designing appropriate reward functions, and state encodings. We provide a family of reward functions, tunable by choice of a parameter. We also investigate the application of different basic and enhanced Q-learning components. Giang Truong, Huu Le, Erchuan Zhang, David Suter, Syed Zulqarnain Gilani |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Maximum Consensus by Weighted Influences of Monotone Boolean FunctionsabstractMaximisation of Consensus (MaxCon) is one of the most widely used robust criteria in computer vision. Tennakoon et al. (CVPR2021), made a connection between MaxCon and estimation of influences of a Monotone Boolean function. In such, there are two distributions involved: the distribution defining the influence measure; and the distribution used for sampling to estimate the influence measure. This paper studies the concept of weighted influences for solving MaxCon. In particular, we study the Bernoulli measures. Theoretically, we prove the weighted influences, under this measure, of points belonging to larger structures are smaller than those of points belonging to smaller structures in general. We also consider another “natural” family of weighting strategies: sampling with uniform measure concentrated on a particular (Hamming) level of the cube. One can choose to have matching distributions: the same for defining the measure as for implementing the sampling. This has the advantage that the sampler is an unbiased estimator of the measure. Based on weighted sampling, we modify the algorithm of Tennakoon et al., and test on both synthetic and real datasets. We show some modest gains of Bernoulli sampling, and we illuminate some of the interactions between structure in data and weighted measures and weighted sampling. Erchuan Zhang, David Suter, Ruwan B. Tennakoon, Tat-Jun Chin, Alireza Bab-Hadiashar, Giang Truong, Syed Zulqarnain Gilani |
CVPR | 7 |
| 2022 | Show, Attend and Detect: Towards Fine-Grained Assessment of Abdominal Aortic Calcification on Vertebral Fracture Assessment Scans
Syed Zulqarnain Gilani, Naeha Sharif, David Suter, John T. Schousboe, Siobhan Reid, William D. Leslie, Joshua R. Lewis |
MICCAI (3) | 1 |
| 2022 | Sparse Hypergraph Community Detection Thresholds in Stochastic Block ModelabstractCommunity detection in random graphs or hypergraphs is an interesting fundamental problem in statistics, machine learning and computer vision. When the hypergraphs are generated by a {\em stochastic block model}, the existence of a sharp threshold on the model parameters for community detection was conjectured by Angelini et al. 2015. In this paper, we confirm the positive part of the conjecture, the possibility of non-trivial reconstruction above the threshold, for the case of two blocks. We do so by comparing the hypergraph stochastic block model with its Erd{\"o}s-R{\'e}nyi counterpart. We also obtain estimates for the parameters of the hypergraph stochastic block model. The methods developed in this paper are generalised from the study of sparse random graphs by Mossel et al. 2015 and are motivated by the work of Yuan et al. 2022. Furthermore, we present some discussion on the negative part of the conjecture, i.e., non-reconstruction of community structures. Erchuan Zhang, David Suter, Giang Truong, Syed Zulqarnain Gilani |
NeurIPS | 4 |
| 2021 | Unsupervised Learning for Robust Fitting: A Reinforcement Learning ApproachabstractRobust model fitting is a core algorithm in a large number of computer vision applications. Solving this problem efficiently for datasets highly contaminated with outliers is, however, still challenging due to the underlying computational complexity. Recent literature has focused on learning-based algorithms. However, most approaches are supervised (which require a large amount of labelled training data). In this paper, we introduce a novel unsupervised learning framework that learns to directly solve robust model fitting. Unlike other methods, our work is agnostic to the underlying input features, and can be easily generalized to a wide variety of LP-type problems with quasi-convex residuals. We empirically show that our method out-performs existing unsupervised learning approaches, and achieves competitive results compared to traditional methods on several important computer vision problems1. Giang Truong, Huu Le, David Suter, Erchuan Zhang, Syed Zulqarnain Gilani |
CVPR | 5 |
| 2021 | Relation Graph Network for 3D Object Detection in Point CloudsabstractConvolutional Neural Networks (CNNs) have emerged as a powerful tool for object detection in 2D images. However, their power has not been fully realised for detecting 3D objects directly in point clouds without conversion to regular grids. Moreover, existing state-of-the-art 3D object detection methods aim to recognize objects individually without exploiting their relationships during learning or inference. In this article, we first propose a strategy that associates the predictions of direction vectors with pseudo geometric centers, leading to a win-win solution for 3D bounding box candidates regression. Secondly, we propose point attention pooling to extract uniform appearance features for each 3D object proposal, benefiting from the learned direction features, semantic features and spatial coordinates of the object points. Finally, the appearance features are used together with the position features to build 3D object-object relationship graphs for all proposals to model their co-existence. We explore the effect of relation graphs on proposals' appearance feature enhancement under supervised and unsupervised settings. The proposed relation graph network comprises a 3D object proposal generation module and a 3D relation module, making it an end-to-end trainable network for detecting 3D objects in point clouds. Experiments on challenging benchmark point cloud datasets (SunRGB-D, ScanNet and KITTI) show that our algorithm performs better than existing state-of-the-art. Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Liang Zhang 0010, Ajmal Mian |
IEEE Trans. Image Process. | 2 |
| 2020 | Point attention network for semantic segmentation of 3D point clouds
Mingtao Feng, Liang Zhang 0010, Xuefei Lin, Syed Zulqarnain Gilani, Ajmal Mian |
Pattern Recognit. | 4 |
| 2019 | Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video CaptioningabstractAutomatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These methods mainly focus on tailoring sequence learning through RNNs for better caption generation, whereas off-the-shelf visual features are borrowed from CNNs. We argue that careful designing of visual features for this task is equally important, and present a visual feature encoding technique to generate semantically rich captions using Gated Recurrent Units (GRUs). Our method embeds rich temporal dynamics in visual features by hierarchically applying Short Fourier Transform to CNN features of the whole video. It additionally derives high level semantics from an object detector to enrich the representation with spatial dynamics of the detected objects. The final representation is projected to a compact space and fed to a language model. By learning a relatively simple language model comprising two GRU layers, we establish new state-of-the-art on MSVD and MSR-VTT datasets for METEOR and ROUGELmetrics. Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Syed Zulqarnain Gilani, Ajmal Mian |
CVPR | 4 |
| 2018 | Learning From Millions of 3D Scans for Large-Scale 3D Face RecognitionabstractDeep networks trained on millions of facial images are believed to be closely approaching human-level performance in face recognition. However, open world face recognition still remains a challenge. Although, 3D face recognition has an inherent edge over its 2D counterpart, it has not benefited from the recent developments in deep learning due to the unavailability of large training as well as large test datasets. Recognition accuracies have already saturated on existing 3D face datasets due to their small gallery sizes. Unlike 2D photographs, 3D facial scans cannot be sourced from the web causing a bottleneck in the development of deep 3D face recognition networks and datasets. In this backdrop, we propose a method for generating a large corpus of labeled 3D face identities and their multiple instances for training and a protocol for merging the most challenging existing 3D datasets for testing. We also propose the first deep CNN model designed specifically for 3D face recognition and trained on 3.1 Million 3D facial scans of 100K identities. Our test dataset comprises 1,853 identities with a single 3D scan in the gallery and another 31K scans as probes, which is several orders of magnitude larger than existing ones. Without fine tuning on this dataset, our network already outperforms state of the art face recognition by over 10%. We fine tune our network on the gallery set to perform end-to-end large scale 3D face recognition which further improves accuracy. Finally, we show the efficacy of our method for the open world face recognition problem. Syed Zulqarnain Gilani, Ajmal Mian |
CVPR | 1 |
| 2018 | 3D Face Reconstruction from Light Field Images: A Model-Free Approach
Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Ajmal Mian |
ECCV (10) | 2 |
| 2018 | Dense 3D Face CorrespondenceabstractWe present an algorithm that automatically establishes dense correspondences between a large number of 3D faces. Starting from automatically detected sparse correspondences on the outer boundary of 3D faces, the algorithm triangulates existing correspondences and expands them iteratively by matching points of distinctive surface curvature along the triangle edges. After exhausting keypoint matches, further correspondences are established by generating evenly distributed points within triangles by evolving level set geodesic curves from the centroids of large triangles. A deformable model (K3DM) is constructed from the dense corresponded faces and an algorithm is proposed for morphing the K3DM to fit unseen faces. This algorithm iterates between rigid alignment of an unseen face followed by regularized morphing of the deformable model. We have extensively evaluated the proposed algorithms on synthetic data and real 3D faces from the FRGCv2, Bosphorus, BU3DFE and UND Ear databases using quantitative and qualitative benchmarks. Our algorithm achieved dense correspondences with a mean localisation error of 1.28 mm on synthetic faces and detected 14 anthropometric landmarks on unseen real faces from the FRGCv2 database with 3 mm precision. Furthermore, our deformable model fitting algorithm achieved 98.5 percent face recognition accuracy on the FRGCv2 and 98.6 percent on Bosphorus database. Our dense model is also able to generalize to unseen datasets. Syed Zulqarnain Gilani, Ajmal Mian, Faisal Shafait, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Deep, dense and accurate 3D face correspondence for generating population specific deformable models
Syed Zulqarnain Gilani, Ajmal Mian, Peter R. Eastwood |
Pattern Recognit. | 1 |
| 2015 | Shape-based automatic detection of a large number of 3D facial landmarksabstractWe present an algorithm for automatic detection of a large number of anthropometric landmarks on 3D faces. Our approach does not use texture and is completely shape based in order to detect landmarks that are morphologically significant. The proposed algorithm evolves level set curves with adaptive geometric speed functions to automatically extract effective seed points for dense correspondence. Correspondences are established by minimizing the bending energy between patches around seed points of given faces to those of a reference face. Given its hierarchical structure, our algorithm is capable of establishing thousands of correspondences between a large number of faces. Finally, a morphable model based on the dense corresponding points is fitted to an unseen query face for transfer of correspondences and hence automatic detection of landmarks. The proposed algorithm can detect any number of pre-defined landmarks including subtle landmarks that are even difficult to detect manually. Extensive experimental comparison on two benchmark databases containing 6, 507 scans shows that our algorithm outperforms six state of the art algorithms. Syed Zulqarnain Gilani, Faisal Shafait, Ajmal Mian |
CVPR | 1 |
| 2014 | Perceptual Differences between Men and Women: A 3D Facial Morphometric PerspectiveabstractUnderstanding the features employed by the human visual system in gender classification is considered a critical step towards improving machine based gender classification systems. We propose the use of 3D Euclidean and geodesic distances between biologically significant facial landmarks to classify gender. We perform five different experiments on the BU-3DFE face database to look for more representative features that can replicate our visual system. Based on our experiments we suggest that the human visual system looks at the ratio of 3D Euclidean and geodesic distance as these features can classify facial gender with an accuracy of 99.32%. The features selected by our proposed gender classification experiment are robust to ethnicity and moderate changes in expression. They also replicate the perceptual gender bias towards certain features and hence become good candidates for being a more representative feature set. Syed Zulqarnain Gilani, Ajmal Mian |
ICPR | 1 |
| 2014 | Gradient based efficient feature selectionabstractSelecting a reduced set of relevant and non-redundant features for supervised classification problems is a challenging task. We propose a gradient based feature selection method which can search the feature space efficiently and select a reduced set of representative features. We test our proposed algorithm on five small and medium sized pattern classification datasets as well as two large 3D face datasets for computer vision applications. Comparison with the state of the art wrapper and filter methods shows that our proposed technique yields better classification results in lesser number of evaluations of the target classifier. The feature subset selected by our algorithm is representative of the classes in the data and has the least variation in classification accuracy. Syed Zulqarnain Gilani, Faisal Shafait, Ajmal Mian |
WACV | 1 |
| 2011 | Data Driven Bandwidth for Medoid Shift Algorithm
Syed Zulqarnain Gilani, Naveed Iqbal Rao |
ICCSA (2) | 1 |
| 2009 | Fast Block Clustering Based Optimized Adaptive Mediod Shift
Syed Zulqarnain Gilani, Naveed Iqbal Rao |
CAIP | 1 |