VLDB 2026 Research / reviewers in the wild / expert
Deepak Mishra 0003
dblp:65/6758-3
· DBLP profile ↗
27ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-4078-9400ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Model Synchronization for Diagnostic Redefinition through a Novel Selective Parameter UnlearningabstractFederated learning (FL) allows multiple medical institutions to collaboratively train machine learning models without sharing sensitive patient data, preserving privacy. However, as medical guidelines and disease classifications change over time, existing models can become outdated and may need updates to stay relevant. We propose a novel approach that efficiently updates federated models by selectively removing outdated knowledge without requiring full retraining. Our approach uses gradient-based Shapley value approximations to identify and modify the most important model parameters linked to obsolete diagnostic categories. This enables precise unlearning of outdated information while preserving performance on current diagnoses. We validate our method on the PathMNIST and COVID-19 Radiography datasets, showing that it can effectively eliminate specific diagnostic classes with minimal loss in accuracy for relevant conditions. Our method only requires a single communication round among clients and offers better control than previous techniques by targeting individual parameters instead of whole convolutional feature channels (in Convolutional Neural Networks (CNN)). This makes it especially useful for keeping federated medical models aligned with evolving medical knowledge.1 Mayank Kumar Kundalwal, Mamta, Deepak Mishra 0003, Asif Ekbal |
WACV | 3 |
| 2025 | Scale-Aware Adaptive Feature Quantization for Robust Medical Image Representation LearningabstractCNNs have become the standard for medical image interpretation, but concerns persist about their reliability in real-world applications. CNNs can be sensitive to small variations in image quality and vulnerable to adversarial attacks, potentially leading to inaccurate diagnoses. To address these issues, we introduce a novel scale-aware adaptive feature quantization approach. This enhances the robustness and reliability of CNNs by adaptively combining quantized representations from multiple scales, improving performance on low-quality or perturbed images. Our approach uses soft codes and dynamic weighting to adaptively combine features from different scales, creating a more informative final quantized representation. Experimental results on diverse medical datasets, including chest X-rays and dermatoscopic images, demonstrate the effectiveness of our approach. Our method significantly outperforms both standard CNNs and state-of-the-art approaches, with substantial gains across all metrics (AUC, F1 score). These improvements range from 2.6% to 11%, demonstrating our method’s superior performance and reliability for medical diagnosis in challenging real-world scenarios. Azad Singh, Deepak Mishra 0003 |
ECAI | 2 |
| 2025 | ViT Coupled Efficient CMOS Image Sensor with Sparse Acquisition and Patch SelectionabstractThis paper addresses the challenges of efficient image acquisition and processing in resource-constrained environments by introducing sparsity-driven CMOS Image Sensor (CIS) architecture coupled with Vision Transformers (ViTs). Our proposed approach incorporates a sensor-level dimensionality reduction technique to capture sparse and high-relevance features for enhancing power and computational efficiency at the sensor level. Key proposals for the pipeline include a re-designed CIS architecture that enables selective feature acquisition through on-sensor edge detection, an adaptive threshold mechanism for reducing ADC operations (> 50% for τ = 0.9) through edge-based pixel selection, and a co-designed memory management strategy focused on patch-wise data retention. Experimental evaluations on CIFAR-10, STL-10, Food-101, and Caltech-256 show a reduction of ≈ 89%, 78%, 76%, and 81% data volume for 100 selected patches while maintaining accuracies of ≈ 94%, 93%, 74%, and 82%, respectively. These improvements in the acquisition architecture demonstrate scalable, real-time processing potential on edge devices, making the proposed architecture a robust solution for low-power applications requiring efficient data acquisition and processing, particularly for ViTs. Wilfred Kisku, Azad Singh, Amandeep Kaur 0005, Deepak Mishra 0003 |
ISCAS | 4 |
| 2025 | Fine-Grained Rib Fracture Diagnosis with Hyperbolic Embeddings: A Detailed Annotation Framework and Multi-label Classification Model
Shripad Pate, Aiman Farooq, Suvrankar Datta, Musadiq Aadil Sheikh, Atin Kumar, Deepak Mishra 0003 |
MICCAI (15) | 6 |
| 2025 | Survival Prediction in Lung Cancer through Multi-Modal Representation LearningabstractSurvival prediction is a crucial task associated with cancer diagnosis and treatment planning. This paper presents a novel approach to survival prediction by harnessing comprehensive informationfrom CT and PET scans, along with associated Genomic data. Current methods rely on either a single modality or the integration of multiple modalities for prediction without adequately addressing associations across patients or modalities. We aim to develop a robust predictive model for survival outcomes by integrating multimodal imaging data with genetic information while accounting for associations across patients and modali-ties. We learn representations for each modality via a self-supervised module and harness the semantic similarities across the patients to ensure the embeddings are aligned closely. However, optimizing solely for global relevance is inadequate, as many pairs sharing similar high-level semantics, such as tumor type, are inadvertently pushed apart in the embedding space. To address this issue, we use a cross-patient module (CPM) designed to harness inter-subject correspondences. The CPM module aims to bring together embeddings from patients with similar disease characteristics. Our experimental evaluation of the dataset of Non-Small Cell Lung Cancer (NSCLC) patients demonstrates the effectiveness of our approach in predicting survival outcomes, outperforming state-of-the-art methods. Aiman Farooq, Deepak Mishra 0003, Santanu Chaudhury |
WACV | 2 |
| 2025 | OTCXR: Rethinking Self-supervised Alignment using Optimal Transport for Chest X-ray AnalysisabstractSelf-supervised learning (SSL) has emerged as a promising technique for analyzing medical modalities such as X-rays due to its ability to learn without annotations. However, conventional SSL methods face challenges in achieving se-mantic alignment and capturing subtle details, which limits their ability to accurately represent the underlying anatom-ical structures and pathological features. To address these limitations, we propose OTCXR, a novel SSL framework that leverages optimal transport (OT) to learn dense seman-tic invariance. By integrating OT with our innovative Cross-Viewpoint Semantics Infusion Module (CV-SIM), OTCXR enhances the model's ability to capture not only local spa-tial features but also global contextual dependencies across different viewpoints. This approach enriches the effective-ness of SSL in the context of chest radiographs. Further-more, OTCXR incorporates variance and covariance reg-ularizations within the OT framework to prioritize clini-cally relevant information while suppressing less informa-tive features. This ensures that the learned representations are comprehensive and discriminative, particularly benefi-cial for tasks such as thoracic disease diagnosis. We vali-date OTCXR's efficacy through comprehensive experiments on three publicly available chest X-ray datasets. Our em-pirical results demonstrate the superiority of OTCXR over state-of-the-art methods across all evaluated tasks, confirming its capability to learn semantically rich representations. Vandan Gorade, Azad Singh, Deepak Mishra 0003 |
WACV | 3 |
| 2025 | Self-Supervised Contextual Representations of Chest X-Ray ImagesabstractSelf-supervised learning (SSL) has gained traction in medical image analysis, enabling representation learning with limited labels. While contrastive SSL, using diverse augmentations, has become dominant, we argue that applying standard augmentations, originally designed for natural images, to medical images like chest X-rays is suboptimal. Chest X-rays possess unique structures and subtle abnormalities that differ from natural images, and preserving these during augmentation is critical for learning clinically meaningful representations. In this paper, we introduce a novel set of domain-specific contextual transformations tailored for chest X-rays, including anatomy-aware perturbations and context random masking, designed to preserve diagnostic semantics during SSL pre-training. This augmentation strategy is the first to explicitly align transformation design with radiological context, addressing a major gap in existing medical SSL approaches. Experiments on NIH, RSNA, and SIIM datasets show that our approach yields up to 5% improvement in downstream tasks under limited supervision, compared to standard augmentations. Azad Singh, Dipan Mandal, Deepak Mishra 0003 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Large Scale Time-Series Representation Learning via Simultaneous Low- and High-Frequency Feature BootstrappingabstractLearning representations from unlabeled time series data is a challenging problem. Most existing self-supervised and unsupervised approaches in the time-series domain fall short in capturing low- and high-frequency features at the same time. As a result, the generalization ability of the learned representations remains limited. Furthermore, some of these methods employ large-scale models like transformers or rely on computationally expensive techniques such as contrastive learning. To tackle these problems, we propose a noncontrastive self-supervised learning (SSL) approach that efficiently captures low- and high-frequency features in a cost-effective manner. The proposed framework comprises a Siamese configuration of a deep neural network with two weight-sharing branches which are followed by low- and high-frequency feature extraction modules. The two branches of the proposed network allow bootstrapping of the latent representation by taking two different augmented views of raw time series data as input. The augmented views are created by applying random transformations sampled from a single set of augmentations. The low- and high-frequency feature extraction modules of the proposed network contain a combination of multilayer perceptron (MLP) and temporal convolutional network (TCN) heads, respectively, which capture the temporal dependencies from the raw input data at various scales due to the varying receptive fields. To demonstrate the robustness of our model, we performed extensive experiments and ablation studies on five real-world time-series datasets. Our method achieves state-of-art performance on all the considered datasets. Vandan Gorade, Azad Singh, Deepak Mishra 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | IDQCE: Instance Discrimination Learning Through Quantized Contextual Embeddings for Medical Images
Azad Singh, Deepak Mishra 0003 |
ICPR (12) | 2 |
| 2024 | Aggregation-Assisted Proxyless Distillation: A Novel Approach for Handling System Heterogeneity in Federated LearningabstractSystem heterogeneity in Federated Learning (FL) is commonly dealt with knowledge distillation by combining the clients’ knowledge via distillation into a global model. However, such knowledge transfer to the global model is often limited by distillation efficiency and unavailability of the client data. Most of the existing approaches require proxy data on the server side for distillation, which often becomes a bottleneck. To circumvent these limitations, we propose a novel FL framework, FedAgPD (Aggregation-Assisted Proxyless Distillation for Heterogeneous Federated Learning) that comprises of deep mutual learning (DML) at client end, and global aggregation followed by noise engineered data-free distillation at the server end. DML enables server side global aggregation which otherwise is infeasible due to different client model architectures. The aggregation results in knowledge integration which is further boosted by the subsequent distillation. We further introduce the idea of selective mutual learning where only those clients perform DML that are not limited by computational capacity. This reduces the overall computational burden without any compromise in the performance. We conduct rigorous experiments on various publicly available datasets and observe a remarkable improvement in the performance over the existing heterogeneous FL methods. For example, for CIFAR100 dataset, FedAgPD shows almost two times better performance as compared to the best baseline. Moreover, we compared FedAgPD with recent homogeneous methods and observed a competitive performance. The results provide evidence for the utility and effectiveness of our approach and open up a new direction for heterogeneous FL. Code for FedAgPD is available at https://github.com/nirbhay-design/FedAgPD Nirbhay Sharma, Mayank Raj, Deepak Mishra 0003 |
IJCNN | 3 |
| 2024 | CoBooM: Codebook Guided Bootstrapping for Medical Image Representation Learning
Azad Singh, Deepak Mishra 0003 |
MICCAI (12) | 2 |
| 2024 | An Intelligent System With Reduced Readout Power and Lightweight CNN for Vision ApplicationsabstractAn always-on intelligent system comprising of an image sensor requires continuous functioning of each pixel. This includes sensing the illumination content of the scene and also the conversion of the analog values into their digital representations. Therefore, power consumption during analog to digital conversion and computational cost at the image sensor module become critical while designing a system that is always-on and incorporates intelligence near the sensor module. This work focuses on the inherent property of the ADC for converting the analog pixel values to digital values by taking a defined number of analog-to-digital converter (ADC) cycles. The design factors considered are 1) Power saving due to reduced ADC conversion cycles for each pixel; 2) The reduced bit-precision of the processing unit to reduce hardware cost; 3) The dataflow design throughhls4ml, which produces parallel computational modes for low latency CNN architectures. The proposed work implements two lightweight CNN models with reduced parameters as compared to the original architectural models of VGG16 (like) and SqueezeNet (like) which are trained in Qkeras and deployed on Zynq UltraScale+ MPSoC board. In addition, the design pipeline is validated on the MobileNetV2 and GhostNet architectures to demonstrate its generalization ability. A detailed analysis shows that limiting the number of ADC bits from 8 to 4 reduces the mean accuracy merely from 50.3 to 49.17 for VGG16 (like) and 67.83 to 67.80 for SqueezeNet (like) model, however, the readout power is significantly reduced from 140.45 mW to 7.7 mW for STL-10 dataset with$96\times96$image resolution. Additional experiments are conducted with CIFAR-10 and mini-ImageNet datasets for classification and with Oxford-IIIT Pet Dataset for segmentation. The proposed work, thus, provides empirical evidence that a reasonable performance for intelligent vision tasks with power saving can be achieved by tuning CNN models to work with reduced ADC bit precision. Wilfred Kisku, Amandeep Kaur 0005, Deepak Mishra 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | MLVICX: Multi-Level Variance-Covariance Exploration for Chest X-Ray Self-Supervised Representation LearningabstractSelf-supervised learning (SSL) reduces the need for manual annotation in deep learning models for medical image analysis. By learning the representations from unablelled data, self-supervised models perform well on tasks that require little to no fine-tuning. However, for medical images, like chest X-rays, characterised by complex anatomical structures and diverse clinical conditions, a need arises for representation learning techniques that encode fine-grained details while preserving the broader contextual information. In this context, we introduce MLVICX (Multi-Level Variance-Covariance Exploration for Chest X-ray Self-Supervised Representation Learning), an approach to capture rich representations in the form of embeddings from chest X-ray images. Central to our approach is a novel multi-level variance and covariance exploration strategy that effectively enables the model to detect diagnostically meaningful patterns while reducing redundancy. MLVICX promotes the retention of critical medical insights by adapting global and local contextual details and enhancing the variance and covariance of the learned embeddings. We demonstrate the performance of MLVICX in advancing self-supervised chest X-ray representation learning through comprehensive experiments. The performance enhancements we observe across various downstream tasks highlight the significance of the proposed approach in enhancing the utility of chest X-ray embeddings for precision medical diagnosis and comprehensive image analysis. For pertaining, we used the NIH-Chest X-ray dataset. Downstream tasks utilized NIH-Chest X-ray, Vinbig-CXR, RSNA pneumonia, and SIIM-ACR Pneumothorax datasets. Overall, we observe up to 3% performance gain over SOTA SSL approaches in various downstream tasks. Additionally, to demonstrate generalizability of our method, we conducted additional experiments on fundus images and observed superior performance on multiple datasets. Codes are available at GitHub. Azad Singh, Vandan Gorade, Deepak Mishra 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Clinically Relevant Myocardium Segmentation in Cardiac Magnetic Resonance ImagesabstractDeep learning approaches have shown great success in myocardium region segmentation in Cardiac MR (CMR) images. However, most of these often ignore irregularities such as protrusions, breaks in contour, etc. As a result, the common practice by clinicians is to manually correct the obtained outputs for the evaluation of myocardium condition. This paper aims to make the deep learning systems capable of handling the aforementioned irregularities and satisfy desired clinical constraints, necessary for various downstream clinical analysis. We propose a refinement model which imposes structural constraints on the outputs of the existing deep learning-based myocardium segmentation methods. The complete system is a pipeline of deep neural networks where an initial network performs myocardium segmentation as accurate as possible and the refinement network removes defects from the initial output to make it suitable for clinical decision support systems. We experiment with datasets collected from four different sources and observe consistent final segmentation outputs with improvement up to 8% in Dice Coefficient and up to 18 pixels in Hausdorff Distance due to the proposed refinement model. The proposed refinement strategy leads to qualitative and quantitative improvements in the performances of all the considered segmentation networks. Our work is an important step towards the development of a fully automatic myocardium segmentation system. It can also be generalized for other tasks where the object of interest has regular structure and the defects can be modelled statistically. Rohit Gavirni, Divij Gupta, Deepak Mishra 0003, Arbind K. Gupta, Sanjaya Viswamitra |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | MBGRLp: Multiscale Bootstrap Graph Representation Learning on Pointcloud (Student Abstract)abstractPoint cloud has gained a lot of attention with the availability of a large amount of point cloud data and increasing applications like city planning and self-driving cars. However, current methods, often rely on labeled information and costly processing, such as converting point cloud to voxel. We propose a self-supervised learning approach to tackle these problems, combating labelling and additional memory cost issues. Our proposed method achieves results comparable to supervised and unsupervised baselines on the widely used benchmark datasets for self-supervised point cloud classification like ShapeNet, ModelNet10/40. Vandan Gorade, Azad Singh, Deepak Mishra 0003 |
AAAI | 3 |
| 2022 | BAFL: Federated Learning with Base Ablation for Cost Effective CommunicationabstractFederated learning is a distributed machine learning setting in which clients train a global model on their local data and share their knowledge with the server in form of the trained model while maintaining privacy of the data. The server aggregates clients' knowledge to create a generalized global model. Two major challenges faced in this process are data heterogeneity and high communication cost. We target the latter and propose a simple approach, BAFL (Federated Learning for Base Ablation) for cost effective communication in federated learning. In contrast to the common practice of employing model compression techniques to reduce the total communication cost, we propose a fine-tuning approach to leverage the feature extraction ability of layers at different depths of deep neural networks. We use a model pretrained on general-purpose large scale data as a global model. This helps in better weight initialization and reduces the total communication cost required for obtaining the generalized model. We achieve further cost reduction by focusing only on the layers responsible for semantic features (data specific information). The clients fine tune only top layers on their local data. Base layers are ablated while transferring the model and clients communicate parameters corresponding to the remaining layers. This results in reduction of communication cost per round without compromising the accuracy. We evaluate the proposed approach using VGG-16 and ResNet-50 models on datasets including WBC, FOOD-101, and CIFAR-10 and obtain up to two orders of reduction in total communication cost as compared to the conventional federated learning. We perform experiments in both IID and Non-IID settings and observe consistent improvements. Mayank Kumar Kundalwal, Anurag Saraswat, Ishan Mishra, Deepak Mishra 0003 |
ICPR | 4 |
| 2021 | On-Array Compressive Acquisition in CMOS Image Sensors Using Accumulated Spatial GradientsabstractA compressive acquisition technique for on-array image compression is proposed in this paper. It capitalizes on representation ability of accumulated spatial gradients of the acquired scene. The local variations inferred from strength of the accumulated gradients are used as cues to vary number of samples read through the image sensor readout. Such sampling enables the reconstruction using traditional interpolation techniques with desired quality. The proposed method is first verified using MATLAB simulations, where on an average, a compression of 87% is achieved, for a threshold of 40 intensity levels. The images are reconstructed using nearest neighbour interpolation (NNI) method which results in a mean peak signal to noise ratio (PSNR) value of 29.09 dB. The reconstructed images are further enhanced using deep convolutional neural network, which improves the PSNR to 32.46 dB. The biggest advantage of the proposed technique is low-complex hardware design. As a proof of concept, a hardware implementation of the technique is performed using discrete components. Pixel intensity values of standard images are converted into analog voltages using a data acquisition system and mapped in the input voltage range of 1.5 V -5.5 V. For a threshold of 3.8 V, the compression of 81% - 83% is observed for the considered images. The proposed technique is simple and effective, and is suitable for low-power complementary metal oxide semiconductor (CMOS) image sensors. Amandeep Kaur 0005, Deepak Mishra 0003, K. M. Amogh, Mukul Sarkar |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Variational Inference with Latent Space Quantization for Adversarial ResilienceabstractDespite their tremendous success in modelling high-dimensional data manifolds, deep neural networks suffer from the threat of adversarial attacks - Existence of perceptually valid input-like samples obtained through careful perturbation that lead to degradation in the performance of the underlying model. Major concerns with existing defense mechanisms include non-generalizability across different attacks, models and large inference time. In this paper, we propose a generalized defense mechanism capitalizing on the expressive power of regularized latent space based generative models. We design an adversarial filter, devoid of access to classifier and adversaries, which makes it usable in tandem with any classifier. The basic idea is to learn a Lipschitz constrained mapping from the data manifold, incorporating adversarial perturbations, to a quantized latent space and re-map it to the true data manifold. Specifically, we simultaneously auto-encode the data manifold and its perturbations implicitly through the perturbations of the regularized and quantized generative latent space, realized using variational inference. We demonstrate the efficacy of the proposed formulation in providing resilience against multiple attack types (black and white box) and methods, while being almost real-time. Our experiments show that the proposed method surpasses the state-of-the-art techniques in several cases. The implementation code is available at - https://github.com/mayank31398/lqvae. Vinay Kyatham, Deepak Mishra 0003, Prathosh A. P. |
ICPR | 2 |
| 2020 | Effect of the Latent Structure on Clustering With GANsabstractGenerative adversarial networks (GANs) have shown remarkable success in the generation of data from natural data manifolds such as images. In several scenarios, it is desirable that generated data is well-clustered, especially when there is severe class imbalance. In this paper, we focus on the problem of clustering in the generated space of GANs and uncover its relationship with the characteristics of the latent space. We derive from first principles, the necessary and sufficient conditions needed to achieve faithful clustering in the GAN framework: (i) presence of a multimodal latent space with adjustable priors, (ii) existence of a latent space inversion mechanism and, (iii) imposition of the desired cluster priors on the latent space. We also identify the GAN models in the literature that partially satisfy these conditions and demonstrate the importance of all the components required, through ablative studies on multiple real-world image datasets. Additionally, we describe a procedure to construct a multimodal latent space which facilitates learning of cluster priors with sparse supervision. Codes for our implementation is available at https://github.com/NEMGAN/NEMGAN-P. Deepak Mishra 0003, Aravind Jayendran, Prathosh A. P. |
IEEE Signal Process. Lett. | 1 |
| 2020 | Target-Independent Domain Adaptation for WBC Classification Using Generative Latent SearchabstractAutomating the classification of camera-obtained microscopic images of White Blood Cells (WBCs) and related cell subtypes has assumed importance since it aids the laborious manual process of review and diagnosis. Several State-Of-The-Art (SOTA) methods developed using Deep Convolutional Neural Networks suffer from the problem of domain shift - severe performance degradation when they are tested on data (target) obtained in a setting different from that of the training (source). The change in the target data might be caused by factors such as differences in camera/microscope types, lenses, lighting-conditions etc. This problem can potentially be solved using Unsupervised Domain Adaptation (UDA) techniques albeit standard algorithms presuppose the existence of a sufficient amount of unlabelled target data which is not always the case with medical images. In this paper, we propose a method for UDA that is devoid of the need for target data. Given a test image from the target data, we obtain its 'closest-clone' from the source data that is used as a proxy in the classifier. We prove the existence of such a clone given that infinite number of data points can be sampled from the source distribution. We propose a method in which a latent-variable generative model based on variational inference is used to simultaneously sample and find the 'closest-clone' from the source distribution through an optimization procedure in the latent space. We demonstrate the efficacy of the proposed method over several SOTA UDA methods for WBC classification on datasets captured using different imaging modalities under multiple settings. Prashant Pandey 0002, Prathosh A. P., Vinay Kyatham, Deepak Mishra 0003, Tathagato Rai Dastidar |
IEEE Trans. Medical Imaging | 4 |
| 2019 | A 12-bit, 2.5-bit/Phase Column-Parallel Cyclic ADCabstractA 12-bit, 1.67-MS/s, two-stage cyclic ADC, using a 1.5-bit algorithm in a 2.5-bit framework is proposed in this brief. The number of accurate comparators is reduced to half as compared with the conventional 2.5-bit stage, which reduces the power consumption. Furthermore, the pipelined operation of the two stages reduces the total number of clock-cycles, which improves the conversion rate. The proposed ADC is designed and fabricated in a standard 180-nm CMOS technology. The obtained differential nonlinearity and integral nonlinearity are +0.5/-0.5 LSB and +0.8/-0.9 LSB, respectively. The ADC consumes 435-μW of power and occupies an area of 0.045 mm2. The postlayout simulations of ADC designed in a column-pitch of 5.6 μm show that it is suitable for column-parallel readout in CMOS image sensors. Amandeep Kaur 0005, Deepak Mishra 0003, Mukul Sarkar |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Despeckling CNN with Ensembles of Classical OutputsabstractUltrasound (US) image despeckling is a problem of high clinical importance. Machine learning solutions to the problem are considered impractical due to the unavailability of speckle-free US image dataset. On the other hand, the classical approaches, which are able to provide the desired outputs, have limitations like input dependent parameter tuning. In this work, a convolutional neural network (CNN) is developed which learns to remove speckle from US images using the outputs of these classical approaches. It is observed that the existing approaches can be combined in a complementary manner to generate an output better than their individual outputs. Thus, the CNN is trained using the individual outputs as well as the output ensembles. It eliminates the cumbersome process of parameter tuning required by the existing approaches for every new input. Further, the proposed CNN is able to outperform the state-of-the-art despeckling approaches and produces the outputs even better than the ensembles for certain images. Deepak Mishra 0003, Sarthak Tyagi, Santanu Chaudhury, Mukul Sarkar, Arvinder Singh Soin |
ICPR | 1 |
| 2018 | Automatic Quantification of Stomata for High-Throughput Plant PhenotypingabstractStomatal morphology is a key phenotypic trait for plants' response analysis under various environmental stresses (e.g. drought, salinity etc.). Stomata exhibit diverse characteristics with respect to orientation, size, shape and varying degree of papillae occlusion. Thus, the biologists currently rely on manual or semi-automatic approaches to accurately compute its morphological traits based on scanning electron microscopic (SEM) images of leaf surface. In contrast to these subjective and low-throughput methods, we propose a novel automated framework for stomata quantification. It is realized based on a hybrid approach where the candidate stomata region is first detected by a convolutional neural network (CNN) and the occlusion is dealt with an inpainting algorithm. In addition, we propose stomata segmentation based quantification framework to solve the problem of shape, scale and occlusion in an end-to-end manner. The performance of the proposed automated frameworks is evaluated by comparing the derived traits with manually computed morphological traits of stomata. With no prior information about its size and location, the hybrid and end-to-end machine learning frameworks shows a correlation of 0.94 and 0.93, respectively on rice stomata images. Furthermore, they successfully enable wheat stomata quantification showing generalizability in terms of cultivars. Swati Bhugra, Deepak Mishra 0003, Anupama Anupama, Santanu Chaudhury, Brejesh Lall, Archana Chugh |
ICPR | 2 |
| 2018 | Segmentation of Vascular Regions in Ultrasound Images: A Deep Learning ApproachabstractVascular region segmentation in ultrasound images is necessary for applications like automatic registration, and surgical navigation. In this paper, a pipelined network comprising of a convolutional neural network (CNN) followed by unsupervised clustering is proposed to perform vessel segmentation in liver ultrasound images. The work is motivated by the tremendous success of CNNs in object detection and localization. CNN here is trained to localize vascular regions, which are subsequently segmented by the clustering. The proposed network results in 99.14% pixel accuracy and 69.62% mean region intersection over union on 132 images. These values are better than some existing methods. Deepak Mishra 0003, Santanu Chaudhury, Mukul Sarkar, Sidharth Manohar, Arvinder Singh Soin |
ISCAS | 1 |
| 2018 | A 12-bit, 2.5-bit/cycle, 1 MS/s two-stage cyclic ADC, for high-speed CMOS Image sensorsabstractA 12-bit, 1 MS/s, two-stage cyclic ADC, with novel 2.5-bit/cycle architecture is proposed in this paper. A 1.5-bit algorithm is used in a 2.5-bit framework, which reduces the required number of accurate comparators and power consumption by 42%. Further, the ADC shows 46% improvement in the conversion rate as compared to the state-of-the-art two-stage cyclic ADC. The proposed ADC is designed and fabricated in a standard 180 nm CMOS technology. The obtained values of DNL and INL are +0.5/-0.5 LSB and +0.8/-0.9 LSB respectively. The ADC consumes 0.8 mW of power and occupies an area of 0.045 mm2with a FoM of 0.19 pJ/conversion-step. The proposed ADC when designed in a column pitch of 5.6 μm, will result in a frame-rate of 1000 frames/sec for a 1 Mpixel array. Amandeep Kaur 0005, Deepak Mishra 0003, Mukul Sarkar |
ISCAS | 2 |
| 2018 | Ultrasound Image Enhancement Using Structure Oriented Adversarial NetworkabstractIn this letter, we aim to develop a deep adversarial despeckling approach to enhance the quality of ultrasound images. Most of the existing approaches target a complete removal of speckle, which produces oversmooth outputs and results in loss of structural details. In contrast, the proposed approach reduces the speckle extent without altering the structural and qualitative attributes of the ultrasound images. A despeckling residual neural network (DRNN) is trained with an adversarial loss imposed by a discriminator. The discriminator tries to differentiate between the despeckled images generated by the DRNN and the set of high-quality images. Further to prevent the developed network from oversmoothing, a structural loss term is used along with the adversarial loss. Experimental evaluations show that the proposed DRNN outperforms the state-of-the-art despeckling approaches in terms of the structural similarity index measure, peak signal to noise ratio, edge preservation index, and speckle region's signal to noise ratio. Deepak Mishra 0003, Santanu Chaudhury, Mukul Sarkar, Arvinder Singh Soin |
IEEE Signal Process. Lett. | 1 |
| 2018 | Edge Probability and Pixel Relativity-Based Speckle Reducing Anisotropic DiffusionabstractAnisotropic diffusion filters are one of the best choices for speckle reduction in the ultrasound images. These filters control the diffusion flux flow using local image statistics and provide the desired speckle suppression. However, inefficient use of edge characteristics results in either oversmooth image or an image containing misinterpreted spurious edges. As a result, the diagnostic quality of the images becomes a concern. To alleviate such problems, a novel anisotropic diffusion-based speckle reducing filter is proposed in this paper. A probability density function of the edges along with pixel relativity information is used to control the diffusion flux flow. The probability density function helps in removing the spurious edges and the pixel relativity reduces the oversmoothing effects. Furthermore, the filtering is performed in superpixel domain to reduce the execution time, wherein a minimum of 15% of the total number of image pixels can be used. For performance evaluation, 31 frames of three synthetic images and 40 real ultrasound images are used. In most of the experiments, the proposed filter shows a better performance as compared to the state-of-the-art filters in terms of the speckle region's signal-to-noise ratio and mean square error. It also shows a comparative performance for figure of merit and structural similarity measure index. Furthermore, in the subjective evaluation, performed by the expert radiologists, the proposed filter's outputs are preferred for the improved contrast and sharpness of the object boundaries. Hence, the proposed filtering framework is suitable to reduce the unwanted speckle and improve the quality of the ultrasound images. Deepak Mishra 0003, Santanu Chaudhury, Mukul Sarkar, Arvinder Singh Soin |
IEEE Trans. Image Process. | 1 |