Basim Azam

dblp:282/7157 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-3367-6467ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unified graph-based framework for visual explainability in convolutional neural networks
abstract
In deep learning, understanding the decision-making processes of complex models is essential for advancing interpretability and trust in artificial intelligence systems. We introduce Causal Relational Attribution Graph (C-RAG), designed to deliver comprehensive, multi-perspective explanations of convolutional neural networks (CNNs) via a graph representation. C-RAG integrates gradient-based local attribution with global feature importance by constructing a graph-based representation that captures hierarchical feature inter-dependencies. In this framework, feature clusters are represented as graph nodes, and their interactions are quantified through combined localized and global attribution metrics, ensuring interpretable insights into model behavior. We evaluate C-RAG across diverse benchmark datasets (ImageNet, CIFAR-10, MNIST) and CNN architectures (ResNet18, VGG19, DenseNet201, LeNet), demonstrating significant advancements over state-of-the-art explainability methods in faithfulness, robustness, and computational efficiency. The proposed approach facilitates accurate spatial feature localization, robust dependency mapping, and efficient explanation generation, making it a valuable tool for critical applications such as medical imaging and autonomous systems. We provide a novel graph-based explainability framework, which bridges the gap between local and global interpretability, C-RAG addresses key limitations in existing methods, establishing a robust foundation for explainable AI in computer vision.
Basim Azam, Thihagoda Gamage Pubudu Sanjeewani, Brijesh K. Verma, Ashfaqur Rahman, Lipo Wang 0001
Inf. Sci.1
2026 A novel neuron efficiency metric for enhancing deep neural network pruning
abstract
Abstract Deep Neural Networks (DNNs) have achieved state-of-the-art performance across various domains, yet their widespread adoption remains constrained by substantial computational and memory demands. While model pruning has emerged as a compelling strategy to address these challenges, existing methods often suffer from critical shortcomings: (1) non-selective neuron pruning which overlooks neuron activation dynamics, leading to the removal of critical neurons; (2) an inability to account for task-specific neuron importance, which causes accuracy degradation; and (3) failure to mitigate redundancy in neuron activations, resulting in suboptimal compression. In this work, we propose a novel Neuron Efficiency Metric (NEM), which integrates three key components—Neuron Activation Rate (NAR), Class-specific Activation Strength (CAS), and Neuron Overlap Index (NOI)—to address these limitations and guide a more effective pruning process. By iteratively evaluating the relevance of each neuron along these axes, NEM ensures selective and structured pruning that minimizes the retention of redundant or irrelevant neurons while preserving task-critical activations. The proposed method is tested on modern architectures using benchmark datasets such as MNIST and CIFAR-10, demonstrating a significant reduction in computational complexity while maintaining or even improving model performance. The results reveal that NEM achieves a higher degree of compression with minimal accuracy loss compared to conventional techniques.
Basim Azam, Brijesh K. Verma, Ashfaqur Rahman, Lipo Wang 0001
Neural Comput. Appl.1
2025 Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control
abstract
Ethical issues around text-to-image (T2I) models demand a comprehensive control over the generative content. Existing techniques addressing these issues for responsible T2I models aim for the generated content to be fair and safe (non-violent/explicit). However, these methods remain bounded to handling the facets of responsibility concepts individually, while also lacking in interpretability. Moreover, they often require alteration to the original model, which compromises the model performance. In this work, we propose a unique technique to enable responsible T2I generation by simultaneously accounting for an extensive range of concepts for fair and safe content generation in a scalable manner. The key idea is to distill the target T2I pipeline with an external plug-and-play mechanism that learns an interpretable composite responsible space for the desired concepts, conditioned on the target T2I pipeline. We use knowledge distillation and concept whitening to enable this. At inference, the learned space is utilized to modulate the generative content. A typical T2I pipeline presents two plug-in points for our approach, namely; the text embedding space and the diffusion model latent space. We develop modules for both points and show the effectiveness of our approach with a range of strong results. Our code can be accessed at https://basim-azam.github.io/responsiblediffusion/
Basim Azam, Naveed Akhtar
CVPR1
2025 GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector
abstract
We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is challenging. Addressing that, we introduce a unique 3D positional information embedded voxel optimization mechanism to fuse multi-view features. To prioritize neural field reconstruction in object regions, we also devise a double importance sampling scheme for the NeRF branch of our detector. We additionally propose an opacity optimization module for precise voxel opacity prediction by enforcing multi-view consistency constraints. Moreover, to further improve voxel density consistency across multiple perspectives, we incorporate ray distance as a weighting factor to minimize cumulative ray errors. Our unique modules synergetically form an end-to-end neural model that establishes new state-of-the-art in NeRF-based multi-view 3D detection, verified with extensive experiments on ScanNet and ARKITScenes. Code will be available at https://github.com/ZechuanLi/GO-N3RDet.
Zechuan Li, Hongshan Yu, Yihao Ding, Jinhao Qiao, Basim Azam, Naveed Akhtar
CVPR5
2025 DDB: Diffusion Driven Balancing to Address Spurious Correlations
Aryan Yazdan Parast, Basim Azam, Naveed Akhtar
ICCV2
2025 A Novel Non-iterative Training Method for CNN Classifiers Using Gram-Schmidt Process
abstract
Abstract Convolutional neural networks have become prominent machine learning models, particularly in the realm of computer vision, due to their ability to predict and extract robust features from raw image data. CNNs, similar to other neural network models, undergo training via backpropagation, an iterative technique. However, the backpropagation algorithm has notable challenges, including slow convergence, susceptibility to local minima, and hypersensitivity to learning rates. These challenges not only impact the model’s accuracy but also make the training process computationally intensive. To address these limitations, We introduce a novel approach that trains the CNN classifier using a non-iterative learning method. The proposed approach involves automatic extraction of pertinent features from the raw-data, followed by the application of Gram–Schmidt process to decompose the feature matrix and determine classifier’s weights. The proposed method has shown enhanced predictive accuracy over state-of-the-art models when evaluated on two benchmark datasets, MNIST and CIFAR-10. The extensive experimentation using most cited pre-trained experiments validate the effectiveness of our proposed method.
Basim Azam, Deepthi Praveenlal Kuttichira, Thihagoda Gamage Pubudu Sanjeewani, Brijesh K. Verma, Ashfaqur Rahman, Lipo Wang 0001
Neural Process. Lett.1
2025 RSVMamba for Tree Species Classification Using UAV RGB Remote Sensing Images
abstract
Effective forest tree species (TS) classification is critical for various application domains such as forest management, biodiversity conservation, and ecological research. However, existing studies on TS classification predominantly rely on high-cost and processing-intensive hyperspectral data, which limits practical applications on large scales. In this work, we focus on investigating the potential of cost-effective unmanned aerial vehicle (UAV) RGB images for TS classification in heterogeneous forests and propose a method that fully leverages the rich spatial, semantic, and visible spectral information of UAV RGB images. We propose an RSVMamba model, which incorporates improved visual state-space (VSS) blocks and an AutoDownsampling module to enhance accuracy and stability while paying particular attention to small objects in sparse spatial locations. The model achieves linear computational complexity while retaining the global receptive field, making it particularly suitable for processing high spatial-resolution images. Additionally, we collected UAV RGB images covering$40~\text {km}^{2}$of subtropical forest in southern China. A meticulous evaluation of this data shows that our method achieves an overall accuracy (OA) of 84.28% for eight TS, dead trees, and other broadleaves. We verify the superiority of our method through a series of comparative experiments on the collected and benchmark datasets. Our results affirm the usefulness of single-temporal UAV RGB images for TS classification in heterogeneous forest environments. Furthermore, the proposed method bridges the gap between data accessibility and precision in TS classification, broadening the boundaries of single-temporal UAV RGB images for practical forestry applications and providing a more cost-effective and time-flexible solution for this problem.
Juntao Gu, Basim Azam, Moule Lin, Chao Li 0066, Weipeng Jing 0001, Naveed Akhtar
IEEE Trans. Geosci. Remote. Sens.3
2024 Optimizing CNNs with Gram Schmidt Non-iterative Learning for Image Recognition
Deepthi Praveenlal Kuttichira, Basim Azam, Brijesh K. Verma, Ashfaqur Rahman, Lipo Wang 0001
ICONIP (3)2
2024 Neuron Efficiency Index: An Empirical Method for Optimizing Parameters in Deep Learning
abstract
Deep Neural Networks (DNNs) have undeniably achieved groundbreaking success across diverse applications. Nevertheless, their complex architectures inherently lead to substantial computational demands and memory prerequisites. To surmount these challenges, this research paper introduces a pioneering approach designed to amplify DNN efficiency via a unique iterative pruning technique Neuron Efficiency Index (NEI), that considers activation frequency of each neuron, class sensitivity and redundancy in the dense layer neurons. The central objective of this method is to curtail the computational burden of the model, all the while ensuring that performance remains intact and enhanced. The proposed technique is used to prune state-of-the-art architectures and comprehensive comparison is presented on benchmark dataset MNIST and CIFAR-10. The evaluation presents that proposed NEI improves the model accuracy while reducing the computations and complexity of the architecture. The work contributes to the field of neural network optimization.
Basim Azam, Deepthi Praveenlal Kuttichira, Brijesh K. Verma
IJCNN1
2023 A Graph-based Context Learning Technique for Image Parsing
abstract
The modern deep learning-based architectures have performed well for pixel-wise segmentation tasks. The consideration of context is of vital importance for generation of accurate semantic information. In this research, a deep learning-based image parsing framework is proposed that utilizes novel relation-aware context learning technique. The proposed technique explores the graph constructs from the training data to learn the co-occurring context associations of object category labels using the graph edge connections. The proposed graph-based context learning technique defines the scene specific relation-awareness among semantic object categories, e.g., the probability of sky, road and building to co-exist in a scene is high. The proposed image parsing architecture (including the novel graph-based context learning technique) is evaluated on the benchmark datasets. In addition, a comprehensive comparison with existing image parsing techniques is presented to establish the efficacy of the scene-graph generation. The in-depth investigation of graph generation is presented to demonstrate the improvement in pixel-wise labeling.
Basim Azam, Brijesh K. Verma
IJCNN1
2022 A Novel Optimized Context-Based Deep Architecture for Scene Parsing
Ranju Mandal, Brijesh K. Verma, Basim Azam, Henry Selvaraj
ICONIP (6)3
2022 Relationship aware context adaptive deep learning for image parsing
Basim Azam, Ranju Mandal, Brijesh K. Verma
Inf. Sci.1
2021 Deep Learning Model with GA-based Visual Feature Selection and Context Integration
abstract
Deep learning models have been very successful in computer vision and image processing applications. Since its inception, Convolutional Neural Network (CNN)-based deep learning models have consistently outperformed other machine learning methods on many significant image processing benchmarks. Many top-performing methods for image segmentation are also based on deep CNN models. However, deep CNN models fail to integrate global and local context alongside visual features despite having complex multi-layer architectures. We propose a novel three-layered deep learning model that learns independently global and local contextual information alongside visual features, and visual feature selection based on a genetic algorithm. The novelty of the proposed model is that One-vs-All binary class-based learners are introduced to learn Genetic Algorithm (GA) optimized features in the visual layer, followed by the contextual layer that learns global and local contexts of an image, and finally the third layer integrates all the information optimally to obtain the final class label. Stanford Background and CamVid benchmark image parsing datasets were used for our model evaluation, and our model shows promising results. The empirical analysis reveals that optimized visual features with global and local contextual information play a significant role to improve accuracy and produce stable predictions comparable to state-of-the-art deep CNN models.
Ranju Mandal, Basim Azam, Brijesh K. Verma, Mengjie Zhang 0001
CEC2
2021 Context-Based Deep Learning Architecture with Optimal Integration Layer for Image Parsing
Ranju Mandal, Basim Azam, Brijesh K. Verma
ICONIP (2)2
2021 Relationship Aware Context Adaptive Feature Selection Framework for Image Parsing
abstract
Feature selection for deep learning architectures is one of the important and challenging steps in developing an efficient image parsing application. In this paper, a novel image parsing architecture which makes use of unique feature selection is proposed. It introduces the idea of weighted relationship awareness to reduce the redundancy of features and optimally select an efficient subset of feature representations. The proposed architecture is evaluated on Cam Vid benchmark dataset. A comparison with state-of-the-art methods was conducted which showed significant improvements in terms of segmentation and classification accuracy.
Basim Azam, Ranju Mandal, Brijesh K. Verma
IJCNN1