Mustansar Fiaz

dblp:12/10812 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0003-2289-2284ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2025 EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
abstract
Automated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and resource management. Existing generic VLMs do not perform well on Remote Sensing data, while the recent Geo-spatial VLMs remain restricted to a fixed resolution and few sensor modalities. In this paper, we introduce EarthDial, a conversational assistant specifically designed for Earth Observation (EO) data, transforming complex, multi-sensory Earth observations into interactive, natural language dialogues. EarthDial supports multi- spectral, multi-temporal, and multi-resolution imagery, enabling a wide range of remote sensing tasks, including classification, detection, captioning, question answering, visual reasoning, and visual grounding. To achieve this, we introduce an extensive instruction tuning dataset comprising over 11.11M instruction pairs covering RGB, Synthetic Aperture Radar (SAR), and multispectral modalities such as Near-Infrared (NIR) and infrared. Furthermore, EarthDial handles bi-temporal and multi-temporal sequence analysis for applications like change detection. Our extensive experimental results on 44 downstream datasets demonstrate that EarthDial outperforms existing generic and domain-specific models, achieving better generalization across various EO tasks. Our source codes and pre-trained models are at https://github.com/hiyamdebary/EarthDial.
Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D. Watson, Levente J. Klein, Fahad Shahbaz Khan, Salman Khan 0001
CVPR4
2025 COSNet: A Novel Semantic Segmentation Network using Enhanced Boundaries in Cluttered Scenes
abstract
Automated waste recycling aims to efficiently separate the recyclable objects from the waste by employing vision-based systems. However, the presence of varying shaped objects having different material types makes it a challenging problem, especially in cluttered environments. Existing segmentation methods perform reasonably on many semantic segmentation datasets by employing multi-contextual representations, however, their performance is degraded when utilized for waste object segmentation in cluttered scenarios. In addition, plastic objects further increase the complexity of the problem due to their translucent nature. To address these limitations, we introduce an efficacious segmentation network, named COSNet, that uses boundary cues along with multi-contextual information to accurately segment the objects in cluttered scenes. COSNet introduces novel components including feature sharpening block (FSB) and boundary enhancement module (BEM) for enhancing the features and highlighting the boundary information of irregular waste objects in cluttered environment. Extensive experiments on three challenging datasets including ZeroWaste-f [4], SpectralWaste [5], and ADE20K [50] demonstrate the effectiveness of the proposed method. Our COSNet achieves a significant gain of 1.8% on ZeroWaste-f and 2.1% on SpectralWaste datasets respectively in terms of mIoU metric. Source code is available at https://github.com/techmn/cosnet.
Mamoona Javaid, Mubashir Noman, Mustansar Fiaz, Salman Khan 0001
WACV4
2024 Fanet: Feature Amplification Network for Semantic Segmentation in Cluttered Background
abstract
Existing deep learning approaches leave out the semantic cues that are crucial in semantic segmentation present in complex scenarios including cluttered backgrounds and translucent objects, etc. To handle these challenges, we propose a feature amplification network (FANet) as a backbone network that incorporates semantic information using a novel feature enhancement module at multi-stages. To achieve this, we propose an adaptive feature enhancement (AFE) block that benefits from both a spatial context module (SCM) and a feature refinement module (FRM) in a parallel fashion. SCM aims to exploit larger kernel leverages for the increased receptive field to handle scale variations in the scene. Whereas our novel FRM is responsible for generating semantic cues that can capture both low-frequency and high-frequency regions for better segmentation tasks. We perform experiments over challenging real-world ZeroWaste-f [1] dataset which contains background-cluttered and translucent objects. Our experimental results demonstrate the state-of-the-art performance compared to existing methods. The source code can be found at https://github.com/techmn/fanet.
Mamoona Javaid, Mubashir Noman, Mustansar Fiaz, Salman Khan 0001
ICIP4
2024 Detection and Characterization of Urban Heat Islands with Machine Learning
abstract
Assessing and understanding the urban scale impacts of extreme climate events is a global necessity. Risks associated with heat, where intra-urban dynamics and rural/urban boundary conditions greatly impact its distribution, are of particular interest as the evolution of climate change and ur-banization persists. Characterizing Urban Heat Island (UHI) effects is dependent on the availability of high-resolution near-surface air temperature maps and a description of the Local Climate Zones (LCZs). This study assesses the applicability of state-of-the-art (SOTA) Artificial Intelligence (AI) techniques for UHI detection and characterization. A Geospatial Foundation Model (GFM) is fine-tuned to predict 2 m air temperature at a 1 km resolution for the urban areas of Johannesburg, South Africa, with mean absolute error measures less than 1.5 °C. UHI characterization is further enabled through a Fully Connected Network (FCN) model for LCZs classification for the same region of interest.
Muaaz Bhamjee, Hiyam Debary, Zaheed Gaffoor, Tamara Govindasamy, Craig Mahlasi, Mustansar Fiaz, Etienne Eben Vos, Levente J. Klein, Sibusisiwe Makhanya, Campbell D. Watson, Julian Kuehnert
IGARSS6
2024 ChangeBind: A Hybrid Change Encoder for Remote Sensing Change Detection
abstract
Change detection (CD) is a fundamental task in remote sensing (RS) which aims to detect the semantic changes between the same geographical regions at different time stamps. Existing convolutional neural networks (CNNs) based approaches often struggle to capture long-range dependencies. Whereas recent transformer-based methods are prone to the dominant global representation and may limit their capabilities to capture the subtle change regions due to the complexity of the objects in the scene. To address these limitations, we propose an effective Siamese-based framework to encode the semantic changes occurring in the bi-temporal RS images. The main focus of our design is to introduce a change encoder that leverages local and global feature representations to capture both subtle and large change feature information from multi-scale features to precisely estimate the change regions. Our experimental study on two challenging CD datasets reveals the merits of our approach and obtains state-of-the-art performance. Code is available at https://github.com/techmn/changebind.
Mubahsir Noman, Mustansar Fiaz, Hisham Cholakkal
IGARSS2
2024 DDAM-PS: Diligent Domain Adaptive Mixer for Person Search
abstract
Person search (PS) is a challenging computer vision problem where the objective is to achieve joint optimization for pedestrian detection and re-identification (ReID). Although previous advancements have shown promising performance in the field under fully and weakly supervised learning fashion, there exists a major gap in investigating the domain adaptation ability of PS models. In this paper, we propose a diligent domain adaptive mixer (DDAM) for person search (DDAP-PS) framework that aims to bridge a gap to improve knowledge transfer from the labeled source domain to the unlabeled target domain. Specifically, we introduce a novel DDAM module that generates moderate mixed-domain representations by combining source and target domain representations. The proposed DDAM module encourages domain mixing to minimize the distance between the two extreme domains, thereby enhancing the ReID task. To achieve this, we introduce two bridge losses and a disparity loss. The objective of the two bridge losses is to guide the moderate mixed-domain representations to maintain an appropriate distance from both the source and target domain representations. The disparity loss aims to prevent the moderate mixed-domain representations from being biased towards either the source or target domains, thereby avoiding overfitting. Furthermore, we address the conflict between the two subtasks, localization and ReID, during domain adaptation. To handle this cross-task conflict, we forcefully decouple the normaware embedding, which aids in better learning of the moderate mixed-domain representation. We conduct experiments to validate the effectiveness of our proposed method. Our approach demonstrates favorable performance on the challenging PRW and CUHK-SYSU datasets. Our source code is publicly available at https://github.com/mustansarfiaz/DDAM-PS.
Mohammed Khaleed Almansoori, Mustansar Fiaz, Hisham Cholakkal
WACV2
2024 Guided-attention and gated-aggregation network for medical image segmentation
Mustansar Fiaz, Mubashir Noman, Hisham Cholakkal, Rao Muhammad Anwer, Jacob Hanna, Fahad Shahbaz Khan
Pattern Recognit.1
2024 ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
abstract
Deep learning has shown remarkable success in remote sensing change detection (CD), aiming to identify semantic change regions between co-registered satellite image pairs acquired at distinct time stamps. However, existing convolutional neural network (CNN) and transformer-based frameworks often struggle to accurately segment semantic change regions. Moreover, transformers-based methods with standard self-attention suffer from quadratic computational complexity with respect to the image resolution, making them less practical for CD tasks with limited training data. To address these issues, we propose an efficient change detection framework, ELGC-Net, which leverages rich contextual information to precisely estimate change regions while reducing the model size. Our ELGC-Net comprises a Siamese encoder, fusion modules, and a decoder. The focus of our design is the introduction of an Efficient Local-Global Context Aggregator (ELGCA) module within the encoder, capturing enhanced global context and local spatial information through a novel pooled-transpose (PT) attention and depthwise convolution, respectively. The PT attention employs pooling operations for robust feature extraction and minimizes computational cost with transposed attention. Extensive experiments on three challenging CD datasets demonstrate that ELGC-Net outperforms existing methods. Compared to the recent transformer-based CD approach (ChangeFormer), ELGC-Net achieves a 1.4% gain in intersection over union (IoU) metric on the LEVIR-CD dataset, while significantly reducing trainable parameters. Our proposed ELGC-Net sets a new state-of-the-art performance in remote sensing change detection benchmarks. Finally, we also introduce ELGC-Net-LW, a lighter variant with significantly reduced computational complexity, suitable for resource-constrained settings, while achieving comparable performance. Our source code is publicly available at https://github.com/techmn/elgcnet.
Mubashir Noman, Mustansar Fiaz, Hisham Cholakkal, Salman Khan 0001, Fahad Shahbaz Khan
IEEE Trans. Geosci. Remote. Sens.2
2024 Remote Sensing Change Detection With Transformers Trained From Scratch
abstract
Current transformer-based change detection (CD) approaches either employ a pre-trained model trained on large-scale image classification ImageNet dataset or rely on first pre-training on another CD dataset and then fine-tuning on the target benchmark. This current strategy is driven by the fact that transformers typically require a large amount of training data to learn inductive biases, which is insufficient in standard CD datasets due to their small size. We develop an end-to-end CD approach with transformers that is trained from scratch and yet achieves state-of-the-art performance on five benchmarks. Instead of using conventional self-attention that struggles to capture inductive biases when trained from scratch, our architecture utilizes a shuffled sparse-attention operation that focuses on selected sparse informative regions to capture the inherent characteristics of the CD data. Moreover, we introduce a change-enhanced feature fusion (CEFF) module to fuse the features from input image pairs by performing a per-channel re-weighting. Our CEFF module aids in enhancing the relevant semantic changes while suppressing the noisy ones. Extensive experiments on five CD datasets reveal the merits of the proposed contributions, achieving gains as high as 1.35% in intersection over union (IoU) score, compared to the best-published results in the literature. Code is available at https://github.com/mustansarfiaz/ScratchFormer.
Mubashir Noman, Mustansar Fiaz, Hisham Cholakkal, Sanath Narayan, Rao Muhammad Anwer, Salman Khan 0001, Fahad Shahbaz Khan
IEEE Trans. Geosci. Remote. Sens.2
2023 SA2-Net: Scale-aware Attention Network for Microscopic Image Segmentation
Mustansar Fiaz, Moein Heidari, Rao Muhammad Anwer, Hisham Cholakkal
BMVC1
2023 PSM-PS: Part-Based Signal Modulation for Person Search
Reem Abdalla Sharif, Mustansar Fiaz, Rao Muhammad Anwer
CAIP (1)2
2023 SAT: Scale-Augmented Transformer for Person Search
abstract
Person search is a challenging computer vision problem where the objective is to simultaneously detect and reidentify a target person from the gallery of whole scene images captured from multiple cameras. Here, the challenges related to underlying detection and re-identification tasks need to be addressed along with a joint optimization of these two tasks. In this paper, we propose a three-stage cascaded Scale-Augmented Transformer (SAT) person search framework. In the three-stage design of our SAT framework, the first stage performs person detection whereas the last two stages performs both detection and re-identification. Considering the contradictory nature of detection and re-identification, in the last two stages, we introduce separate norm feature embeddings for the two tasks to reconcile the relationship between them in a joint person search model. Our SAT framework benefits from the attributes of convolutional neural networks and transformers by introducing a convolutional encoder and a scale modulator within each stage. Here, the convolutional encoder increases the generalization ability of the model whereas the scale modulator performs context aggregation at different granularity levels to aid in handling pose/scale variations within a region of interest. To further improve the performance during occlusion, we apply shifting augmentation operations at each granularity level within the scale modulator. Experimental results on challenging CUHK-SYSU [35] and PRW [47] datasets demonstrate the favorable performance of our method compared to state-of-the-art methods. Our source code and trained models are available at this https URL.
Mustansar Fiaz, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan
WACV1
2022 PS-ARM: An End-to-End Attention-Aware Relation Mixer Network for Person Search
Mustansar Fiaz, Hisham Cholakkal, Sanath Narayan, Rao Muhammad Anwer, Fahad Shahbaz Khan
ACCV (5)1
2021 4G-VOS: Video Object Segmentation using guided context embedding
Mustansar Fiaz, Muhammad Zaigham Zaheer, Arif Mahmood, Seung-Ik Lee, Soon Ki Jung
Knowl. Based Syst.1