Mannudeep K. Kalra

dblp:64/7612 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0001-9938-7476ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Phrase-Grounded Fact-Checking for Automatically Generated Chest X-Ray Reports
Razi Mahmood, Diego Machado Reyes, Joy T. Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep K. Kalra, Ge Wang 0001, Pingkun Yan, Tanveer F. Syeda-Mahmood
MICCAI (7)7
2025 Navigating the artificial intelligence revolution in neuro-oncology: A multidisciplinary viewpoint
Sanjay Saxena, Soumyaranjan Panda, Ekta Tiwari, Mostafa Fouda, Mannudeep K. Kalra, Ketan Kotecha, Luca Saba, Jasjit S. Suri
Neurocomputing6
2025 Chest X-Ray Foundation Model With Global and Local Representations Integration
abstract
Chest X-ray (CXR) is the most frequently ordered imaging test, supporting diverse clinical tasks from thoracic disease detection to postoperative monitoring. However, task-specific classification models are limited in scope, require costly labeled data, and lack generalizability to out-of-distribution datasets. To address these challenges, we introduce CheXFound, a self-supervised vision foundation model that learns robust CXR representations and generalizes effectively across a wide range of downstream tasks. We pretrained CheXFound on a curated CXR-987K dataset, comprising over approximately 987K unique CXRs from 12 publicly available sources. We propose a Global and Local Representations Integration (GLoRI) head for downstream adaptations, by incorporating fine- and coarse-grained disease-specific local features with global image features for enhanced performance in multilabel classification. Our experimental results showed that CheXFound outperformed state-of-the-art models in classifying 40 disease findings across different prevalence levels on the CXR-LT 24 dataset and exhibited superior label efficiency on downstream tasks with limited training data. Additionally, CheXFound achieved significant improvements on downstream tasks with out-of-distribution datasets, including opportunistic cardiovascular disease risk estimation, mortality prediction, malpositioned tube detection, and anatomical structure segmentation. The above results demonstrate CheXFound's strong generalization capabilities, which will enable diverse downstream adaptations with improved label efficiency in future applications. The project source code is publicly available at https://github.com/RPIDIAL/CheXFound.
Zefan Yang, Xuanang Xu, Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE Trans. Medical Imaging5
2025 Disease-Informed Adaptation of Vision-Language Models
abstract
Expertise scarcity and high cost of data annotation hinder the development of artificial intelligence (AI) foundation models for medical image analysis. Transfer learning provides a way to utilize the off-the-shelf foundation models to address the clinical challenges. However, such models encounter difficulties when adapting to new diseases not presented in their original pre-training datasets. Compounding this challenge is the limited availability of example cases for a new disease, which further leads to the poor performance of the existing transfer learning techniques. This paper proposes a novel method for transfer learning of foundation Vision-Language Models (VLMs) to efficiently adapt them to a new disease with only a few examples. Such an effective adaptation of VLMs hinges on learning the nuanced representation of new disease concepts. By capitalizing on the joint visual-linguistic capabilities of VLMs, we introduce disease-informed contextual prompting in a novel disease prototype learning framework, which enables VLMs to quickly grasp the concept of the new disease, even with limited data. Extensive experiments across multiple pre-trained medical VLMs and multiple tasks showcase the notable enhancements in performance compared to other existing adaptation techniques. The code will be made publicly available at https://github.com/RPIDIAL/Disease-informed-VLM-Adaptation.
Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE Trans. Medical Imaging3
2024 Cardiovascular Disease Detection from Multi-view Chest X-Rays with BI-Mamba
Zefan Yang, Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (5)4
2024 Disease-Informed Adaptation of Vision-Language Models
Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (11)3
2022 Overlooked Trustworthiness of Saliency Maps
Jiajin Zhang, Hanqing Chao, Giridhar Dasegowda, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (3)5
2021 Integrative analysis for COVID-19 patient outcome prediction
Hanqing Chao, Xi Fang 0002, Jiajin Zhang, Fatemeh Homayounieh, Chiara Daniela Arru, Subba R. Digumarthy, Rosa Babaei, Hadi Karimi Mobin, Iman Mohseni, Luca Saba, Alessandro Carriero, Zeno Falaschi, Alessio Pasche, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
Medical Image Anal.15
2021 Quantifying and leveraging predictive uncertainty for medical image assessment
Florin C. Ghesu, Bogdan Georgescu, Awais Mansoor, Youngjin Yoo, Eli Gibson, R. S. Vishwanath, Abishek Balachandran, James M. Balter, Subba R. Digumarthy, Mannudeep K. Kalra, Sasa Grbic, Dorin Comaniciu
Medical Image Anal.12
2021 Deep metric learning-based image retrieval system for chest radiograph and its clinical applications in COVID-19
Aoxiao Zhong, Xiang Li 0001, Dufan Wu, Hui Ren 0001, Kyung Sang Kim, Young-Gon Kim, Varun Buch, Nir Neumark, Bernardo Bizzo, Won Young Tak, Soo Young Park, Yu Rim Lee, Min Kyu Kang, Jung Gil Park, Byung Seok Kim, Woo Jin Chung, Ittai Dayan, Mannudeep K. Kalra, Quanzheng Li
Medical Image Anal.19
2021 Deep Interactive Denoiser (DID) for X-Ray Computed Tomography
abstract
Low-dose computed tomography (LDCT) is desirable for both diagnostic imaging and image-guided interventions. Denoisers are widely used to improve the quality of LDCT. Deep learning (DL)-based denoisers have shown state-of-the-art performance and are becoming mainstream methods. However, there are two challenges to using DL-based denoisers: 1) a trained model typically does not generate different image candidates with different noise-resolution tradeoffs, which are sometimes needed for different clinical tasks; and 2) the model's generalizability might be an issue when the noise level in the testing images differs from that in the training dataset. To address these two challenges, in this work, we introduce a lightweight optimization process that can run on top of any existing DL-based denoiser during the testing phase to generate multiple image candidates with different noise-resolution tradeoffs suitable for different clinical tasks in real time. Consequently, our method allows users to interact with the denoiser to efficiently review various image candidates and quickly pick the desired one; thus, we termed this method deep interactive denoiser (DID). Experimental results demonstrated that DID can deliver multiple image candidates with different noise-resolution tradeoffs and shows great generalizability across various network architectures, as well as training and testing datasets with various noise levels.
Ti Bai, Biling Wang, Dan Nguyen, Bao Wang 0001, Bin Dong 0001, Wenxiang Cong, Mannudeep K. Kalra, Steve B. Jiang
IEEE Trans. Medical Imaging7
2020 Shape and margin-aware lung nodule classification in low-dose CT images via soft activation mapping
Yukun Tian, Hongming Shan, Junping Zhang, Ge Wang 0001, Mannudeep K. Kalra
Medical Image Anal.6
2020 Knowledge-Based Analysis for Mortality Prediction From CT Images
abstract
Low-Dose CT (LDCT) can significantly improve the accuracy of lung cancer diagnosis and thus reduce cancer deaths compared to chest X-ray. The lung cancer risk population is also at high risk of other deadly diseases, for instance, cardiovascular diseases. Therefore, predicting the all-cause mortality risks of this population is of great importance. This paper introduces a knowledge-based analytical method using deep convolutional neural network (CNN) for all-cause mortality prediction. The underlying approach combines structural image features extracted from CNNs, based on LDCT volume at different scales, and clinical knowledge obtained from quantitative measurements, to predict the mortality risk of lung cancer screening subjects. The proposed method is referred as Knowledge-based Analysis of Mortality Prediction Network (KAMP-Net). It constitutes a collaborative framework that utilizes both imaging features and anatomical information, instead of completely relying on automatic feature extraction. Our work demonstrates the feasibility of incorporating quantitative clinical measurements to assist CNNs in all-cause mortality prediction from chest LDCT images. The results of this study confirm that radiologist defined features can complement CNNs in performance improvement. The experiments demonstrate that KAMP-Net can achieve a superior performance when compared to other methods.
Hengtao Guo, Uwe Krüger 0001, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE J. Biomed. Health Informatics4
2020 Severity and Consolidation Quantification of COVID-19 From CT Images Using Deep Learning Based on Hybrid Weak Labels
abstract
Early and accurate diagnosis of Coronavirus disease (COVID-19) is essential for patient isolation and contact tracing so that the spread of infection can be limited. Computed tomography (CT) can provide important information in COVID-19, especially for patients with moderate to severe disease as well as those with worsening cardiopulmonary status. As an automatic tool, deep learning methods can be utilized to perform semantic segmentation of affected lung regions, which is important to establish disease severity and prognosis prediction. Both the extent and type of pulmonary opacities help assess disease severity. However, manually pixel-level multi-class labelling is time-consuming, subjective, and non-quantitative. In this article, we proposed a hybrid weak label-based deep learning method that utilize both the manually annotated pulmonary opacities from COVID-19 pneumonia and the patient-level disease-type information available from the clinical report. A UNet was firstly trained with semantic labels to segment the total infected region. It was used to initialize another UNet, which was trained to segment the consolidations with patient-level information using the Expectation-Maximization (EM) algorithm. To demonstrate the performance of the proposed method, multi-institutional CT datasets from Iran, Italy, South Korea, and the United States were utilized. Results show that our proposed method can predict the infected regions as well as the consolidation regions with good correlation to human annotation.
Dufan Wu, Kuang Gong, Chiara Daniela Arru, Fatemeh Homayounieh, Bernardo Bizzo, Varun Buch, Hui Ren 0001, Kyung Sang Kim, Nir Neumark, Nuobei Xie, Won Young Tak, Soo Young Park, Yu Rim Lee, Min Kyu Kang, Jung Gil Park, Alessandro Carriero, Luca Saba, Mahsa Masjedi, Hamidreza Talari, Rosa Babaei, Hadi Karimi Mobin, Shadi Ebrahimian, Ittai Dayan, Mannudeep K. Kalra, Quanzheng Li
IEEE J. Biomed. Health Informatics27
2020 Quadratic Autoencoder (Q-AE) for Low-Dose CT Denoising
abstract
Inspired by complexity and diversity of biological neurons, our group proposed quadratic neurons by replacing the inner product in current artificial neurons with a quadratic operation on input data, thereby enhancing the capability of an individual neuron. Along this direction, we are motivated to evaluate the power of quadratic neurons in popular network architectures, simulating human-like learning in the form of "quadratic-neuron-based deep learning". Our prior theoretical studies have shown important merits of quadratic neurons and networks in representation, efficiency, and interpretability. In this paper, we use quadratic neurons to construct an encoder-decoder structure, referred as the quadratic autoencoder, and apply it to low-dose CT denoising. The experimental results on the Mayo low-dose CT dataset demonstrate the utility and robustness of quadratic autoencoder in terms of image denoising and model efficiency. To our best knowledge, this is the first time that the deep learning approach is implemented with a new type of neurons and demonstrates a significant potential in the medical imaging field.
Fenglei Fan, Hongming Shan, Mannudeep K. Kalra, Guhan Qian, Matthew Getzin, Yueyang Teng, Juergen Hahn, Ge Wang 0001
IEEE Trans. Medical Imaging3
2019 Quantifying and Leveraging Classification Uncertainty for Chest Radiograph Assessment
Florin C. Ghesu, Bogdan Georgescu, Eli Gibson, Sebastian Gündel, Mannudeep K. Kalra, Subba R. Digumarthy, Sasa Grbic, Dorin Comaniciu
MICCAI (6)5
2018 3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained Network
abstract
Low-dose computed tomography (LDCT) has attracted major attention in the medical imaging field, since CT-associated X-ray radiation carries health risks for patients. The reduction of the CT radiation dose, however, compromises the signal-to-noise ratio, which affects image quality and diagnostic performance. Recently, deep-learning-based algorithms have achieved promising results in LDCT denoising, especially convolutional neural network (CNN) and generative adversarial network (GAN) architectures. This paper introduces a conveying path-based convolutional encoder-decoder (CPCE) network in 2-D and 3-D configurations within the GAN framework for LDCT denoising. A novel feature of this approach is that an initial 3-D CPCE denoising model can be directly obtained by extending a trained 2-D CNN, which is then fine-tuned to incorporate 3-D spatial information from adjacent slices. Based on the transfer learning from 2-D to 3-D, the 3-D network converges faster and achieves a better denoising performance when compared with a training from scratch. By comparing the CPCE network with recently published work based on the simulated Mayo data set and the real MGH data set, we demonstrate that the 3-D CPCE denoising model has a better performance in that it suppresses image noise and preserves subtle structures.
Hongming Shan, Yi Zhang 0018, Qingsong Yang, Uwe Krüger 0001, Mannudeep K. Kalra, Ling Sun 0006, Wenxiang Cong, Ge Wang 0001
IEEE Trans. Medical Imaging5
2018 Correction for "3D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2D Trained Network"
abstract
In[1], please note the updated figure captions for Figures 5, 6, 7, and 8 as follows:
Hongming Shan, Yi Zhang 0018, Qingsong Yang, Uwe Krüger 0001, Mannudeep K. Kalra, Ling Sun 0006, Wenxiang Cong, Ge Wang 0001
IEEE Trans. Medical Imaging5
2018 Low-Dose CT Image Denoising Using a Generative Adversarial Network With Wasserstein Distance and Perceptual Loss
abstract
The continuous development and extensive use of computed tomography (CT) in medical practice has raised a public concern over the associated radiation dose to the patient. Reducing the radiation dose may lead to increased noise and artifacts, which can adversely affect the radiologists' judgment and confidence. Hence, advanced image reconstruction from low-dose CT data is needed to improve the diagnostic performance, which is a challenging problem due to its ill-posed nature. Over the past years, various low-dose CT methods have produced impressive results. However, most of the algorithms developed for this application, including the recently popularized deep learning techniques, aim for minimizing the mean-squared error (MSE) between a denoised CT image and the ground truth under generic penalties. Although the peak signal-to-noise ratio is improved, MSE- or weighted-MSE-based methods can compromise the visibility of important structural details after aggressive denoising. This paper introduces a new CT image denoising method based on the generative adversarial network (GAN) with Wasserstein distance and perceptual similarity. The Wasserstein distance is a key concept of the optimal transport theory and promises to improve the performance of GAN. The perceptual loss suppresses noise by comparing the perceptual features of a denoised output against those of the ground truth in an established feature space, while the GAN focuses more on migrating the data noise distribution from strong to weak statistically. Therefore, our proposed method transfers our knowledge of visual perception to the image denoising task and is capable of not only reducing the image noise level but also trying to keep the critical information at the same time. Promising results have been obtained in our experiments with clinical CT images.
Qingsong Yang, Pingkun Yan, Hengyong Yu, Yongyi Shi, Xuanqin Mou, Mannudeep K. Kalra, Yi Zhang 0018, Ling Sun 0006, Ge Wang 0001
IEEE Trans. Medical Imaging7
2017 Low-Dose CT With a Residual Encoder-Decoder Convolutional Neural Network
abstract
Given the potential risk of X-ray radiation to the patient, low-dose CT has attracted a considerable interest in the medical imaging field. Currently, the main stream low-dose CT methods include vendor-specific sinogram domain filtration and iterative reconstruction algorithms, but they need to access raw data, whose formats are not transparent to most users. Due to the difficulty of modeling the statistical characteristics in the image domain, the existing methods for directly processing reconstructed images cannot eliminate image noise very well while keeping structural details. Inspired by the idea of deep learning, here we combine the autoencoder, deconvolution network, and shortcut connections into the residual encoder-decoder convolutional neural network (RED-CNN) for low-dose CT imaging. After patch-based training, the proposed RED-CNN achieves a competitive performance relative to the-state-of-art methods in both simulated and clinical cases. Especially, our method has been favorably evaluated in terms of noise suppression, structural preservation, and lesion detection.
Hu Chen 0002, Yi Zhang 0018, Mannudeep K. Kalra, Feng Lin 0010, Yang Chen 0008, Peixi Liao, Jiliu Zhou, Ge Wang 0001
IEEE Trans. Medical Imaging3
2017 Comparison Between Pre-Log and Post-Log Statistical Models in Ultra-Low-Dose CT Reconstruction
abstract
X-ray detectors in clinical computed tomography (CT) usually operate in current-integrating mode. Their complicated signal statistics often lead to intractable likelihood functions for practical use in model-based image reconstruction (MBIR). It is therefore desirable to design simplified statistical models without losing the essential factors. Depending on whether the CT transmission data are logarithmically transformed, pre-log and post-log models are two major categories of choices in CT MBIR. Both being approximations, it remains an open question whether one model can notably improve image quality over the other on real scanners. In this study, we develop and compare several pre-log and post-log MBIR algorithms under a unified framework. Their reconstruction accuracy based on simulation and clinical datasets are evaluated. The results show that pre-log MBIR can achieve notably better quantitative accuracy than post-log MBIR in ultra-low-dose CT, although in less extreme cases, post-log MBIR with handcrafted pre-processing remains a competitive alternative. Pre-log MBIR could play a growing role in emerging ultra-low-dose CT applications.
Tzu-Cheng Lee, Soo Mee Kim, Adam M. Alessio, Paul E. Kinahan, Zhiqian Chang, Ken D. Sauer, Mannudeep K. Kalra, Bruno De Man
IEEE Trans. Medical Imaging8