Sivaramakrishnan Rajaraman

dblp:149/5858 · also Sivarama Krishnan Rajaraman · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-0871-8634ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating hallucinations in synthesized clinical texts to improve multimodal deep learning for dermatology
abstract
• An investigation into the effects of pairing synthesized clinical notes with image data to train a multimodal AI algorithm, using dermatology as example problem domain. • Leveraging metadata information to drive clinical note synthesis reduces hallucinations in Large Language Model (LLM) outputs. • Clinical notes generated by different LLMs using metadata lead to similar performance on downstream tasks when paired with real dermatology images. • Combination of multimodal data improves generalization performance on external datasets. Despite recent advancements in the development of foundation models and multimodal (MM) architectures in dermatology, their translation to clinical practice remains limited by the scarcity of large-scale multimodal (MM) datasets, as most publicly available resources are small, unimodal, and lack expressive clinical text. This paper investigates strategies to synthesize and exploit clinical notes paired with dermatological images to effectively train a MM architecture, focusing on solutions to limit the inherently hallucinated contents introduced by Large Language Models (LLMs) and to identify conditions under which synthetic clinical notes can be reliably leveraged. The paper proposes a MM architecture trained on real dermatological images paired with LLM-synthesized clinical notes. We systematically evaluate different note generation strategies, including metadata-guided prompting, alignment of image representations with specific keywords, sentence-level filtering of clinical notes, network architectural designs. Experiments involve 16,000 image-note couples collected from six public datasets for model training and over 37,000 images from fifteen public datasets as external data for generalization assessment. Performance is assessed on cross-modal retrieval and zero-shot learning tasks to quantify robustness and generalization. Results show that metadata inclusion into the prompts reduces the hallucinations within LLM outputs, providing more reliable notes. The resulting MM model trained with these notes show superior performance on multiple downstream tasks. Synthesized clinical notes can be paired with real dermatology images under specific conditions, providing a valuable resource to develop foundation models that can help reduce the dermatologists’ workload.
Niccolò Marini, Zhaohui Liang, Sivaramakrishnan Rajaraman, Zhiyun Xue, Sameer K. Antani
J. Biomed. Informatics3
2025 The Hidden Threat of Hallucinations in Binary Chest X-Ray Pneumonia Classification
abstract
Hallucination in deep learning (DL) classification, where DL models yield confidently erroneous predictions remains a pressing concern. This study investigates whether binary classifiers are truly learning disease-specific features when distinguishing overlapping radiological presentations among pneumonia subtypes on chest X-ray (CXR) images. Specifically, we evaluate if uncertainty measure is a valuable tool in classifying signs of different pathogen-specific subtypes of pneumonia. We evaluated two binary classifiers to classify bacterial pneumonia and viral pneumonia, respectively, from normal CXRs. A third classifier explored the ability to distinguish bacterial from viral pneumonia presentation to highlight our concern regarding the observed hallucinations in the former cases. Our comprehensive analysis computes the Matthews Correlation Coefficient and prediction entropy metrics on a pediatric CXR dataset and reveals that the normal/bacterial and normal/viral classifiers consistently and confidently misclassify the unseen pneumonia subtype to their respective disease class. These findings expose a critical limitation concerning the tendency of binary classifiers to hallucinate by relying on general pneumonia indicators rather than pathogen-specific patterns, thereby challenging their utility in clinical workflows.
Sivaramakrishnan Rajaraman, Zhaohui Liang, Niccolò Marini, Zhiyun Xue, Sameer K. Antani
CBMS1
2024 Addressing Class Imbalance with Latent Diffusion-based Data Augmentation for Improving Disease Classification in Pediatric Chest X-rays
abstract
Deep learning (DL) has transformed medical image classification; however, its efficacy is often limited by significant data imbalance due to far fewer cases (minority class) compared to controls (majority class). It has been shown that synthetic image augmentation techniques can simulate clinical variability, leading to enhanced model performance. We hypothesize that they could also mitigate the challenge of data imbalance, thereby addressing overfitting to the majority class and enhancing generalization. Recently, latent diffusion models (LDMs) have shown promise in synthesizing high-quality medical images. This study evaluates the effectiveness of a text-guided image-to-image LDM in synthesizing disease-positive chest X-rays (CXRs) and augmenting a pediatric CXR dataset to improve classification performance. We first establish baseline performance by fine-tuning an ImageNet-pretrained Inception-V3 model on class-imbalanced data for two tasks-normal vs. pneumonia and normal vs. bronchopneumonia. Next, we fine-tune individual text-guided image-to-image LDMs to generate CXRs showing signs of pneumonia and bronchopneumonia. The Inception-V3 model is retrained on an updated data set that includes these synthesized images as part of augmented training and validation sets. Classification performance is compared using balanced accuracy, sensitivity, specificity, F-score, Matthews correlation coefficient (MCC), Kappa, and Youden's index against the baseline performance. Results show that the augmentation significantly improves Youden's index (p<0.05) and markedly enhances other metrics, indicating that data augmentation using LDM-synthesized images is an effective strategy for addressing class imbalance in medical image classification.
Sivaramakrishnan Rajaraman, Zhaohui Liang, Zhiyun Xue, Sameer K. Antani
BIBM1
2023 Emergency Department Wait Time Forecast based on Semantic and Time Series Patterns in COVID-19 Pandemic
abstract
This study introduces a new ensemble architecture to improve the wait time forecast for healthcare service in the emergency department (ED) of hospital. The new model first used a fine-tuned text embedding model to extract the contextual semantic meaning of patients’ chief complaint from the electronic patient records to estimate the degree of case urgency and combined to a recurrent neural network to process the regular ED wait time patterns. Four text embedding models including the universal sentence encoder with DAN and transformer encoders, the NNLM, and the Swivel were used for semantic analysis. The results show that the new ensemble model can reduce the prediction errors maximumly by 20.0% in mean of absolute error (MAE), 46.0% in mean of squared error (MSE), and 26.6% in root mean squared error (RMSE). A 5-fold cross validation verified that the new model is robust to the ED wait time prediction before and during the COVID-19 pandemic. We conclude that the new model provides an innovative approach to apply semantic analysis of natural language processing to the domain of time series prediction in the healthcare domain.
Zhaohui Liang, Zhiyun Xue, Sivaramakrishnan Rajaraman, Jimmy Huang 0001, Sameer K. Antani
BIBM4
2023 A Study on Reducing Big Data Image Annotation Burden Through Iterative Expert-In-The-Loop Strategy
abstract
A key challenge in development of reliable and robust medical imaging machine learning solution is the lack of annotated data. This problem becomes particularly significant when big data sets are used. These pose a burden on the annotators to manually segment regions of interest which is a labor intensive and tedious approach. One solution toward addressing this challenge is to use an iterative expert-in-the-loop approach where models that are initially, albeit weakly, trained on a small expert segmented data set are progressively used to expand the training data. In this work, we explore the viability of this approach through two segmentation experiments. The first is a challenging problem of segmenting the buccal mucosa region from photographs of the mouth for subsequent detection and classification of lesions aimed at an oral cancer prediction application. The other is to segment the lung region in chest X-ray (CXR) images. For simplicity and to focus on discovering viability and any associated shortcomings, we limited our scope to just using an off-the-shelf U-Net algorithm to determine if this approach to training data expansion improved segmentation results. Our findings show that for the buccal mucosa segmentation in oral photographs, the method achieved up to 10% improvement in Dice Similarity Coefficient over three iterations on a blinded manually segmented hold-out test set before the performance plateaued. However, the training data set size almost doubled in size in two iterations. For CXR lung segmentation, we observe slight performance improvement (1% in one iteration) of the method from its initial model which already has a much higher performance (93%). We analyze the performance of the approach for these data sets and comment on the potential of a human expert-in-the-loop method for training data expansion for unlabeled or weakly-labeled medical imaging data.
Evanjelin Mahmoodi, Zhiyun Xue, Sivaramakrishnan Rajaraman, Sameer K. Antani
BIBM3
2023 Can deep adult lung segmentation models generalize to the pediatric population?
abstract
Lung segmentation in chest X-rays (CXRs) is an important prerequisite for improving the specificity of diagnoses of cardiopulmonary diseases in a clinical decision support system. Current deep learning models for lung segmentation are trained and evaluated on CXR datasets in which the radiographic projections are captured predominantly from the adult population. However, the shape of the lungs is reported to be significantly different across the developmental stages from infancy to adulthood. This might result in age-related data domain shifts that would adversely impact lung segmentation performance when the models trained on the adult population are deployed for pediatric lung segmentation. In this work, our goal is to (i) analyze the generalizability of deep adult lung segmentation models to the pediatric population and (ii) improve performance through a stage-wise, systematic approach consisting of CXR modality-specific weight initializations, stacked ensembles, and an ensemble of stacked ensembles. To evaluate segmentation performance and generalizability, novel evaluation metrics consisting of mean lung contour distance (MLCD) and average hash score (AHS) are proposed in addition to the multi-scale structural similarity index measure (MS-SSIM), the intersection of union (IoU), Dice score, 95% Hausdorff distance (HD95), and average symmetric surface distance (ASSD). Our results showed a significant improvement (p < 0.05) in cross-domain generalization through our approach. This study could serve as a paradigm to analyze the cross-domain generalizability of deep segmentation models for other medical imaging modalities and applications.
Sivaramakrishnan Rajaraman, Feng Yang 0010, Ghada Zamzmi, Zhiyun Xue, Sameer K. Antani
Expert Syst. Appl.1
2022 Real-time echocardiography image analysis and quantification of cardiac indices
abstract
Deep learning has a huge potential to transform echocardiography in clinical practice and point of care ultrasound testing by providing real-time analysis of cardiac structure and function. Automated echocardiography analysis is benefited through use of machine learning for tasks such as image quality assessment, view classification, cardiac region segmentation, and quantification of diagnostic indices. By taking advantage of high-performing deep neural networks, we propose a novel and eicient real-time system for echocardiography analysis and quantification. Our system uses a self-supervised modality-specific representation trained using a publicly available large-scale dataset. The trained representation is used to enhance the learning of target echo tasks with relatively small datasets. We also present a novel Trilateral Attention Network (TaNet) for real-time cardiac region segmentation. The proposed network uses a module for region localization and three lightweight pathways for encoding rich low-level, textural, and high-level features. Feature embeddings from these individual pathways are then aggregated for cardiac region segmentation. This network is fine-tuned using a joint loss function and training strategy. We extensively evaluate the proposed system and its components, which are echo view retrieval, cardiac segmentation, and quantification, using four echocardiography datasets. Our experimental results show a consistent improvement in the performance of echocardiography analysis tasks with enhanced computational eiciency that charts a path toward its adoption in clinical practice. Specifically, our results show superior real-time performance in retrieving good quality echo from individual cardiac view, segmenting cardiac chambers with complex overlaps, and extracting cardiac indices that highly agree with the experts' values. The source code of our implementation can be found in the project's GitHub page.
Ghada Zamzmi, Sivaramakrishnan Rajaraman, Li-Yueh Hsu, Vandana Sachdev, Sameer K. Antani
Medical Image Anal.2
2018 Gender Detection from Spine X-Ray Images Using Deep Learning
abstract
The algorithm described in this paper aims to classify the spine x-ray images according to image characteristics that exhibit gender. We developed a customized sequential CNN model which is trained from scratch using the spine images first and tested it on the NHANES II dataset hosted by the U.S. National Library of Medicine (NLM). Aiming to improve the performance, we then developed a method for extracting the region-of-interest (ROI) in the cervical spine images using a content-based image retrieval (CBIR) method and compared the results of using the original images vs. the ROI images. Later, we applied/tested the method of fine-tuning a DenseNet model that was pre-trained with the ImageNet dataset with the spine images, and this approach gets the best result, achieving classification accuracy of 99% for cervical spine image set and 98% for the lumbar spine image set.
Zhiyun Xue, Sivaramakrishnan Rajaraman, L. Rodney Long, Sameer K. Antani, George R. Thoma
CBMS2