Azad Singh

dblp:319/2836 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6607-1130ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scale-Aware Adaptive Feature Quantization for Robust Medical Image Representation Learning
abstract
CNNs have become the standard for medical image interpretation, but concerns persist about their reliability in real-world applications. CNNs can be sensitive to small variations in image quality and vulnerable to adversarial attacks, potentially leading to inaccurate diagnoses. To address these issues, we introduce a novel scale-aware adaptive feature quantization approach. This enhances the robustness and reliability of CNNs by adaptively combining quantized representations from multiple scales, improving performance on low-quality or perturbed images. Our approach uses soft codes and dynamic weighting to adaptively combine features from different scales, creating a more informative final quantized representation. Experimental results on diverse medical datasets, including chest X-rays and dermatoscopic images, demonstrate the effectiveness of our approach. Our method significantly outperforms both standard CNNs and state-of-the-art approaches, with substantial gains across all metrics (AUC, F1 score). These improvements range from 2.6% to 11%, demonstrating our method’s superior performance and reliability for medical diagnosis in challenging real-world scenarios.
Azad Singh, Deepak Mishra 0003
ECAI1
2025 ViT Coupled Efficient CMOS Image Sensor with Sparse Acquisition and Patch Selection
abstract
This paper addresses the challenges of efficient image acquisition and processing in resource-constrained environments by introducing sparsity-driven CMOS Image Sensor (CIS) architecture coupled with Vision Transformers (ViTs). Our proposed approach incorporates a sensor-level dimensionality reduction technique to capture sparse and high-relevance features for enhancing power and computational efficiency at the sensor level. Key proposals for the pipeline include a re-designed CIS architecture that enables selective feature acquisition through on-sensor edge detection, an adaptive threshold mechanism for reducing ADC operations (> 50% for τ = 0.9) through edge-based pixel selection, and a co-designed memory management strategy focused on patch-wise data retention. Experimental evaluations on CIFAR-10, STL-10, Food-101, and Caltech-256 show a reduction of ≈ 89%, 78%, 76%, and 81% data volume for 100 selected patches while maintaining accuracies of ≈ 94%, 93%, 74%, and 82%, respectively. These improvements in the acquisition architecture demonstrate scalable, real-time processing potential on edge devices, making the proposed architecture a robust solution for low-power applications requiring efficient data acquisition and processing, particularly for ViTs.
Wilfred Kisku, Azad Singh, Amandeep Kaur 0005, Deepak Mishra 0003
ISCAS2
2025 OTCXR: Rethinking Self-supervised Alignment using Optimal Transport for Chest X-ray Analysis
abstract
Self-supervised learning (SSL) has emerged as a promising technique for analyzing medical modalities such as X-rays due to its ability to learn without annotations. However, conventional SSL methods face challenges in achieving se-mantic alignment and capturing subtle details, which limits their ability to accurately represent the underlying anatom-ical structures and pathological features. To address these limitations, we propose OTCXR, a novel SSL framework that leverages optimal transport (OT) to learn dense seman-tic invariance. By integrating OT with our innovative Cross-Viewpoint Semantics Infusion Module (CV-SIM), OTCXR enhances the model's ability to capture not only local spa-tial features but also global contextual dependencies across different viewpoints. This approach enriches the effective-ness of SSL in the context of chest radiographs. Further-more, OTCXR incorporates variance and covariance reg-ularizations within the OT framework to prioritize clini-cally relevant information while suppressing less informa-tive features. This ensures that the learned representations are comprehensive and discriminative, particularly benefi-cial for tasks such as thoracic disease diagnosis. We vali-date OTCXR's efficacy through comprehensive experiments on three publicly available chest X-ray datasets. Our em-pirical results demonstrate the superiority of OTCXR over state-of-the-art methods across all evaluated tasks, confirming its capability to learn semantically rich representations.
Vandan Gorade, Azad Singh, Deepak Mishra 0003
WACV2
2025 Self-Supervised Contextual Representations of Chest X-Ray Images
abstract
Self-supervised learning (SSL) has gained traction in medical image analysis, enabling representation learning with limited labels. While contrastive SSL, using diverse augmentations, has become dominant, we argue that applying standard augmentations, originally designed for natural images, to medical images like chest X-rays is suboptimal. Chest X-rays possess unique structures and subtle abnormalities that differ from natural images, and preserving these during augmentation is critical for learning clinically meaningful representations. In this paper, we introduce a novel set of domain-specific contextual transformations tailored for chest X-rays, including anatomy-aware perturbations and context random masking, designed to preserve diagnostic semantics during SSL pre-training. This augmentation strategy is the first to explicitly align transformation design with radiological context, addressing a major gap in existing medical SSL approaches. Experiments on NIH, RSNA, and SIIM datasets show that our approach yields up to 5% improvement in downstream tasks under limited supervision, compared to standard augmentations.
Azad Singh, Dipan Mandal, Deepak Mishra 0003
IEEE Signal Process. Lett.1
2025 Large Scale Time-Series Representation Learning via Simultaneous Low- and High-Frequency Feature Bootstrapping
abstract
Learning representations from unlabeled time series data is a challenging problem. Most existing self-supervised and unsupervised approaches in the time-series domain fall short in capturing low- and high-frequency features at the same time. As a result, the generalization ability of the learned representations remains limited. Furthermore, some of these methods employ large-scale models like transformers or rely on computationally expensive techniques such as contrastive learning. To tackle these problems, we propose a noncontrastive self-supervised learning (SSL) approach that efficiently captures low- and high-frequency features in a cost-effective manner. The proposed framework comprises a Siamese configuration of a deep neural network with two weight-sharing branches which are followed by low- and high-frequency feature extraction modules. The two branches of the proposed network allow bootstrapping of the latent representation by taking two different augmented views of raw time series data as input. The augmented views are created by applying random transformations sampled from a single set of augmentations. The low- and high-frequency feature extraction modules of the proposed network contain a combination of multilayer perceptron (MLP) and temporal convolutional network (TCN) heads, respectively, which capture the temporal dependencies from the raw input data at various scales due to the varying receptive fields. To demonstrate the robustness of our model, we performed extensive experiments and ablation studies on five real-world time-series datasets. Our method achieves state-of-art performance on all the considered datasets.
Vandan Gorade, Azad Singh, Deepak Mishra 0003
IEEE Trans. Neural Networks Learn. Syst.2
2024 IDQCE: Instance Discrimination Learning Through Quantized Contextual Embeddings for Medical Images
Azad Singh, Deepak Mishra 0003
ICPR (12)1
2024 CoBooM: Codebook Guided Bootstrapping for Medical Image Representation Learning
Azad Singh, Deepak Mishra 0003
MICCAI (12)1
2024 MLVICX: Multi-Level Variance-Covariance Exploration for Chest X-Ray Self-Supervised Representation Learning
abstract
Self-supervised learning (SSL) reduces the need for manual annotation in deep learning models for medical image analysis. By learning the representations from unablelled data, self-supervised models perform well on tasks that require little to no fine-tuning. However, for medical images, like chest X-rays, characterised by complex anatomical structures and diverse clinical conditions, a need arises for representation learning techniques that encode fine-grained details while preserving the broader contextual information. In this context, we introduce MLVICX (Multi-Level Variance-Covariance Exploration for Chest X-ray Self-Supervised Representation Learning), an approach to capture rich representations in the form of embeddings from chest X-ray images. Central to our approach is a novel multi-level variance and covariance exploration strategy that effectively enables the model to detect diagnostically meaningful patterns while reducing redundancy. MLVICX promotes the retention of critical medical insights by adapting global and local contextual details and enhancing the variance and covariance of the learned embeddings. We demonstrate the performance of MLVICX in advancing self-supervised chest X-ray representation learning through comprehensive experiments. The performance enhancements we observe across various downstream tasks highlight the significance of the proposed approach in enhancing the utility of chest X-ray embeddings for precision medical diagnosis and comprehensive image analysis. For pertaining, we used the NIH-Chest X-ray dataset. Downstream tasks utilized NIH-Chest X-ray, Vinbig-CXR, RSNA pneumonia, and SIIM-ACR Pneumothorax datasets. Overall, we observe up to 3% performance gain over SOTA SSL approaches in various downstream tasks. Additionally, to demonstrate generalizability of our method, we conducted additional experiments on fundus images and observed superior performance on multiple datasets. Codes are available at GitHub.
Azad Singh, Vandan Gorade, Deepak Mishra 0003
IEEE J. Biomed. Health Informatics1
2022 MBGRLp: Multiscale Bootstrap Graph Representation Learning on Pointcloud (Student Abstract)
abstract
Point cloud has gained a lot of attention with the availability of a large amount of point cloud data and increasing applications like city planning and self-driving cars. However, current methods, often rely on labeled information and costly processing, such as converting point cloud to voxel. We propose a self-supervised learning approach to tackle these problems, combating labelling and additional memory cost issues. Our proposed method achieves results comparable to supervised and unsupervised baselines on the widely used benchmark datasets for self-supervised point cloud classification like ShapeNet, ModelNet10/40.
Vandan Gorade, Azad Singh, Deepak Mishra 0003
AAAI2