Anwaar Ulhaq

dblp:219/5727 · also Anwaar Ul Haq, Anwaar Ul-Haq · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-5145-7276ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Soft-Masked Transformer for Point Cloud Processing With Skip Attention-Based Upsampling
abstract
Point cloud processing methods leverage local and global point features to cater to downstream tasks, yet they often overlook the task-level context inherent in point clouds during the encoding stage. We argue that integrating task-level information into the encoding stage significantly enhances performance. To that end, we propose SMTransformer which incorporates task-level information into a vector-based transformer by utilizing a soft mask generated from task-level queries and keys to learn the attention weights. Additionally, to facilitate effective communication between features from the encoding and decoding layers in high-level tasks such as segmentation, we introduce a skip-attention-based up-sampling block. This block dynamically fuses features from various resolution points across the encoding and decoding layers. To mitigate the increase in network parameters and training time resulting from the complexity of the aforementioned blocks, we propose a novel shared point position encoding strategy. This strategy allows various transformer blocks to share the same position information over the same resolution points, thereby reducing network parameters and training time without compromising accuracy. Experimental comparisons with existing methods on multiple datasets demonstrate the efficacy of SMTransformer and skip-attention-based up-sampling for semantic segmentation task. In particular, we achieve state-of-the-art semantic segmentation results of 73.9% mIoU on S3DIS Area 5 and 62.4% mIoU on SWAN dataset. Note to Practitioners—Point cloud processing underpins automation tasks such as robotic perception, navigation, and inspection, where accurate 3D understanding is essential. Existing methods often prioritize vision benchmarks while overlooking automation needs like efficiency on limited hardware and robustness in real-world environments. The proposed SMTransformer embeds task-level guidance into feature learning and employs skip-attention up-sampling to improve segmentation accuracy with practical efficiency. It is well-suited for robotic manipulation, autonomous driving, and inspection applications. Current limitations include reliance on GPUs and sensitivity to extreme density variations. Future work will target edge-device deployment and multi-task extensions.
Yong He 0012, Hongshan Yu, Chaoxu Mu, Mingtao Feng, Tongjia Chen, Zechuan Li, Anwaar Ulhaq, Ajmal Mian
IEEE Trans Autom. Sci. Eng.7
2025 Policy Gradient-Based Optimal Subset Selection for Few-Shot Vision-Language Learning
abstract
Vision-Language models (VLMs) like Contrastive Language-Image Pre-Training (CLIP) have been extensively adapted for few-shot classification. Most few-shot methods rely on randomly selected samples from the dataset. However, since only a few samples are used, the sample selection process can significantly impact the performance of the downstream classification task. In this work, we propose a reinforcement learning-based policy gradient technique that employs a diversity and informativeness-based reward function to optimise the sample selection process. We evaluate various sample selection techniques based on downstream classification accuracy across three benchmark datasets, where the proposed method demonstrates promising results.
Muhammad Khizer Ali, Manoranjan Paul, Anwaar Ulhaq, Muhammad Haris Khan, Quazi Mamun
ICIP3
2025 Organoid-ICLIP: Class Imbalance-Aware Vision-Language Learning for Organoid Mitosis Classification
abstract
Visual analysis of brain organoid imaging is critical for advancing the understanding of in vitro organoid development. However, precise differentiation of mitotic stages remains challenging due to complex cellular morphology and limited imaging data. To address these limitations, we present Organoid-ICLIP, a novel Class Imbalance-Aware Contrastive Language-Image Pretraining model tailored for mitotic nuclei analysis. This approach integrates functional and morphological descriptions of mitotic phases by aligning organoid microscopy images with textual descriptions during visual pattern extraction. The proposed approach introduces a new Mitosis Phase Imbalance-Aware Balancing Loss (MPIA-BL), which dynamically adjusts instance weighting based on mitotic phase distribution to address severe class imbalance. By explicitly addressing both inter-class and intra-class imbalance, Organoid-ICLIP achieves more than 6% improvement over traditional contrastive learning approaches on evaluation metrics. Experimental evaluations on the benchmark Borg brain organoid imaging dataset demonstrate the enhanced performance of the proposed approach in mitotic phase classification compared to SOTA vision models.
Anabia Sohail, Oishi Deb, Anwaar Ulhaq
ICIP3
2025 Video anomaly detection in 10 years: a survey and outlook
Moshira Abdalla, Sajid Javed, Muaz Al Radi, Anwaar Ulhaq, Naoufel Werghi
Neural Comput. Appl.4
2024 OrphaGPT: An Adapted Large Language Model for Orphan Diseases Classification
Kushal Pokhrel, Cesar Sanín, Md. Rafiqul Islam 0004, Md. Kowsar Hossain Sakib, Anwaar Ulhaq, Edward Szczerbicki
ACIIDS (1)5
2023 SIDVis: Designing Visual Interactive System for Analyzing Suicide Ideation Detection
abstract
Suicide is a critical global issue that demands a comprehensive examination of factors such as mental illness, substance abuse, financial stress, and trauma. Effectively identifying individuals at risk is vital for intervention and prevention efforts. However, distinguishing suicidal ideation (SID) from non-suicidal language poses challenges. Existing research has addressed this issue, but limited attention has been given to visually interpretable and interactive systems tailored for SID. This study contributes to responsible AI by leveraging deep learning and machine learning techniques to enhance SID detection, enabling proactive interventions and support. In this paper, we introduce SIDVis, an interactive visualization system that improves performance and interpretability at the same time. The rigorous evaluation demonstrates that SIDVis not only outperforms existing methods in terms of accuracy but also provides an explanation for the responsible use of the underlying AI approach, demonstrating its potential to improve SID detection and intervention strategies.
Md. Rafiqul Islam 0004, Md. Kowsar Hossain Sakib, Anwaar Ulhaq, Shanjita Akter, Jianlong Zhou, David Asirvatham
IV3
2022 Dynamic Point Cloud Compression with Cross-Sectional Approach
Faranak Tohidi, Manoranjan Paul, Anwaar Ulhaq
PSIVT3
2021 Features Of ICU Admission In X-Ray Images Of Covid-19 Patients
abstract
This paper presents an original methodology for extracting semantic features from X-rays images that correlate to severity from a data set with patient ICU admission labels through interpretable models. The validation is partially performed by a proposed method that correlates the extracted features with a separate larger data set that does not contain the ICU-outcome labels. The analysis points out that a few features explain most of the variance between patients admitted in ICUs or not. The methods herein can be viewed as a statistical approach highlighting the importance of features related to ICU admission that may have been only qualitatively reported. In between features shown to be over-represented in the external data set were ones like ‘Consolidation’ (1.67), ‘Alveolar’ (1.33), and ‘Effusion’ (1.3). A brief analysis on the locations also showed higher frequency in labels like ‘Bilateral’ (1.58) and Peripheral (1.28) in patients labelled with higher chances to be admitted in ICU. To properly handle the limited data sets, a state-of-the-art lung segmentation network was also trained and presented, together with the use of low-complexity and interpretable models to avoid overfitting.
Douglas P. S. Gomes, Anwaar Ulhaq, Manoranjan Paul, Michael J. Horry, Subrata Chakraborty, Manash Saha, Tanmoy Debnath, D. M. Motiur Rahaman
ICIP2
2020 Active contours with local and global energy based-on fuzzy clustering and maximum a posterior probability for retinal vessel detection
abstract
Summary The performance of active contour model is limited on retinal vessel segmentation as vessel images are usually corrupted with intensity inhomogeneity, low contrast, and weak boundary, which severely affect the segmentation results of retinal vessels. A new active contour model combining the local and global information is proposed in this paper to facilitate the vessel segmentation. In our model, the fuzzy conception is firstly introduced as fuzzy methods generally provide more accurate and robust clustering and the concept of fuzziness in fuzzy clustering, which is represented by membership, can reflect the intensity distribution of the image. Then, we define local energy based on Maximum a Posterior Probability and use spatially varying parameters, mean and stand deviation, to describe the local Gaussian distribution in order to better deal with intensity inhomogeneity. Furthermore, we combine local and global energy based on fuzzy clustering, with a weight coefficient. The coefficient is computed by a weight function according to contrast ratio of the image. Experiments on synthetic and real images and comparisons with other state‐of‐the‐art active contour models show that the proposed model can detect objects more accurate and robust, especially for vessels on retinal angiogram.
Xiancheng Wang, Zhangwei Jiang, Roozbeh Zarei, Guangyan Huang, Anwaar Ulhaq, Xiaoxia Yin, Mengjiao Guo, Jing He 0004
Concurr. Comput. Pract. Exp.6
2019 A Framework for Early Detection of Antisocial Behavior on Twitter Using Natural Language Processing
Ravinder Singh, Jiahua Du, Yanchun Zhang, Hua Wang 0002, Yuan Miao 0001, Omid Ameri Sianaki, Anwaar Ulhaq
CISIS7
2019 Multi-temporal Registration of Environmental Imagery Using Affine Invariant Convolutional Features
Asim Khan, Anwaar Ulhaq, Randall W. Robinson
PSIVT2
2018 Action Recognition in the Dark via Deep Representation Learning
abstract
Human action recognition for automated video surveillance applications is an interesting but a daunting task especially if the videos are captured in unfavourable lighting conditions. These situations encourage the use of multi-sensor video streams. However, simultaneous activity recognition from multiple video streams is a difficult problem due to their complementary and noisy nature. This paper proposes simultaneous action recognition from multiple video streams using deep multi-view representation learning. Furthermore, it introduces a spatio-temporal feature based correlation filter, for simultaneous detection and recognition of multiple human actions in low-light conditions. We evaluated the performance of our proposed filter with extensive experimentation on nighttime action datasets. Experimental results indicate the effectiveness of deep fusion scheme for robust action recognition in extremely low-light conditions.
Anwaar Ulhaq
IPAS1
2018 Deep Cross-view Convolutional Features for View-invariant Action Recognition
abstract
Convolutional neural network (CNN) based approaches have proved very effective for recognizing actions from a fixed viewpoint. However, these approaches are not generalized for recognizing actions captured from arbitrary viewpoint. In this paper, we present a deep multi-view framework for cross-view action recognition. We integrate spatiotemporal convolutional features from multiple views using deep multi-view representation learning. It helps to extract deep discriminative cross-view convolutional features for action recognition from any arbitrary viewpoint. To speed-up action detection and recognition, we then, train a feature based correlation filter for each action class. The proposed framework helps to recognize actions across different view-points with increased accuracy. An extensive experimentation to evaluate the underlying design on four publicly available datasets indicates that problem of view variations in in a single action class can be solved by learning discriminative information from multiple view.
Anwaar Ulhaq
IPAS1
2018 On Space-Time Filtering Framework for Matching Human Actions Across Different Viewpoints
abstract
Space-time template matching is considered as a promising approach for human action recognition. However, a major drawback of template-based methods is computational overhead due to matching in spatial domain. Recently, space-time correlation-based action filters have been proposed for recognizing human actions in frequency domain. These action filters present reduction in time complexity as Fourier transform-based matching is faster than spatial template matching. However, the utility of such action filters is challenged due to a number of factors: 1) inability to deal with view variations due to implicit lack of support for view-invariance; 2) these filters can be trained only for one action class at a time, and separate filters are required for each action class with increased computational overhead; 3) these filters simply take average of similar action instances and behave no better than average filters; and 4) slightly misaligned action data sets create problems as these filters are not shift-invariant. In this paper, we try to address these shortcomings by proposing an advanced space-time filtering framework for recognizing human actions despite large viewpoint variations. Rather than using crude intensity values, we use 3D tensor structure at each pixel, which characterizes the most common local motion in action sequences. Discrete tensor Fourier transform is then applied to achieve frequency domain representations. Then, we form view clusters from multiple view action data and use space-time correlation filtering to achieve discriminative view representations. These representations are used in an innovative way to achieve action recognition despite viewpoint variations. Extensive experimentation is performed on well-known multiple view action data sets, including IXMAS, WVU, and N-UCLA action data set. A detailed performance comparison with the existing view-invariant action recognition techniques indicates that our approach works equally well for RGB and RGB-D video data with increased accuracy and efficiency.
Anwaar Ulhaq, Xiao-Xia Yin, Jing He 0004, Yanchun Zhang
IEEE Trans. Image Process.1
2016 Action-02MCF: A Robust Space-Time Correlation Filter for Action Recognition in Clutter and Adverse Lighting Conditions
Anwaar Ulhaq, Xiao-Xia Yin, Yunchan Zhang, Iqbal Gondal
ACIVS1
2016 FACE: Fully Automated Context Enhancement for night-time video sequences
Anwaar Ulhaq, Xiao-Xia Yin, Jing He 0004, Yanchun Zhang
J. Vis. Commun. Image Represent.1
2013 On Temporal Order Invariance for View-Invariant Action Recognition
abstract
View-invariant action recognition is one of the most challenging problems in computer vision. Various representations are being devised for matching actions across different viewpoints to achieve view invariance. In this paper, we explore the invariance property of temporal order of action instances during action execution and utilize it for devising a new view-invariant action recognition approach. To ensure temporal order during matching, we utilize spatiotemporal features, feature fusion and temporal order consistency constraint. We start by extracting spatiotemporal cuboid features from video sequences and applying feature fusion to encapsulate within-class similarity for the same viewpoints. For each action class, we construct a feature fusion table to facilitate feature matching across different views. An action matching score is then calculated based on global temporal order constraint and number of matching features. Finally, the action label of the class with the maximum value of the matching score is assigned to the query action. Experimentation is performed on multiple view Inria Xmas motion acquisition sequences and West Virginia University action datasets, with encouraging results, that are comparable to the existing view-invariant action recognition techniques.
Anwaar Ulhaq, Iqbal Gondal, M. Manzur Murshed
IEEE Trans. Circuits Syst. Video Technol.1
2011 On dynamic scene geometry for view-invariant action matching
abstract
Variation in viewpoints poses significant challenges to action recognition. One popular way of encoding view-invariant action representation is based on the exploitation of epipolar geometry between different views of the same action. Majority of representative work considers detection of landmark points and their tracking by assuming that motion trajectories for all landmark points on human body are available throughout the course of an action. Unfortunately, due to occlusion and noise, detection and tracking of these landmarks is not always robust. To facilitate it, some of the work assumes that such trajectories are manually marked which is a clear drawback and lacks automation introduced by computer vision. In this paper, we address this problem by proposing view invariant action matching score based on epipolar geometry between actor silhouettes, without tracking and explicit point correspondences. In addition, we explore multi-body epipolar constraint which facilitates to work on original action volumes without any pre-processing. We show that multi-body fundamental matrix captures the geometry of dynamic action scenes and helps devising an action matching score across different views without any prior segmentation of actors. Extensive experimentation on challenging view invariant action datasets shows that our approach not only removes long standing assumptions but also achieves significant improvement in recognition accuracy and retrieval.
Anwaar Ulhaq, Iqbal Gondal, M. Manzur Murshed
CVPR1
2010 A novel color image fusion QoS measure for multi-sensor night vision applications
abstract
Color image fusion of visible and infra-red imagery can play an important role in multi-sensor night vision systems that are an integral part of modern warfare. Image fusion minimizes the amount of required bandwidth by transmitting the fused image rather than multiple sensor images. Color image fusion can be achieved by combining inputs from original colored sensors or by employing pseudo colorization and color transfer to grayscale images. Various quality measures have been proposed for multi-sensor grayscale image fusion techniques; but no appropriate quality measure has been devised for the quality evaluation of multi-sensor color image fusion. In this paper, we propose a novel color image fusion quality measure, Color Fusion Objective Index (CFOI) based on colorfulness, gradient similarity and mutual information techniques. Experimental results show the effectiveness of CFOI to evaluate the color and salient feature extraction introduced by color fusion techniques into the final fused imagery as well as its consistency with subjective evaluation.
Anwaar Ulhaq, Iqbal Gondal, M. Manzur Murshed
ISCC1
2010 Automated multi-sensor color video fusion for nighttime video surveillance
abstract
In this paper, we present an automated color transfer based video fusion method to attain real-time color night vision capability for night-time video surveillance. We utilize simple RGB Color transfer technique to fused pseudo colored video frames without conversion to any uncorrelated color space. We investigated that final color fusion results greatly depend on the selection of target color Image. Therefore, rather than using any arbitrary target color image based on mere general visual anticipation, we have automated target color image selection using structural similarity and color saturation. We further apply color enhancement to improve final appearance of color fused images. Subjective and objective quality evaluations greatly indicate the effectiveness of our color video fusion method for nighttime video surveillance applications.
Anwaar Ulhaq, Iqbal Gondal, M. Manzur Murshed
ISCC1