Pranav Rajpurkar

dblp:161/3744 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-8030-3727ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 36% Language models and text generation · 23% Question answering and dialogue systems · 16%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Medical and health informatics · 100%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model › domain-specific vision-language model
medical vision-language model
0.912025
FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models · CVPR 2025
Computer vision › Vision and language › medical report generation
radiology report generation
0.912025
FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models · CVPR 2025
Medical and health informatics › medical report generation
radiology report generation
0.912025
FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models · CVPR 2025
Machine learning › Deep learning architectures and training
foundation model
0.712023
Multimodal Clinical Benchmark for Emergency Care (MC-BEC): A Comprehensive Benchmark for Evaluating Foundation Models in Emergency Medicine · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.712023
Exploring the Boundaries of GPT-4 in Radiology · EMNLP 2023
Medical and health informatics
clinical prediction
0.712023
Multimodal Clinical Benchmark for Emergency Care (MC-BEC): A Comprehensive Benchmark for Evaluating Foundation Models in Emergency Medicine · NeurIPS 2023
Medical and health informatics
clinical text processing
0.712023
Exploring the Boundaries of GPT-4 in Radiology · EMNLP 2023
Medical and health informatics › radiology › radiology informatics
radiology report analysis
0.712023
Exploring the Boundaries of GPT-4 in Radiology · EMNLP 2023
Computer vision › Image recognition and object detection › medical image analysis
medical image classification
0.412019
CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison · AAAI 2019
Medical and health informatics
medical imaging
0.412019
CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison · AAAI 2019
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
extractive question answering
0.212016
SQuAD: 100, 000+ Questions for Machine Comprehension of Text · EMNLP 2016
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.212016
SQuAD: 100, 000+ Questions for Machine Comprehension of Text · EMNLP 2016
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
reading comprehension datasets
0.212016
SQuAD: 100, 000+ Questions for Machine Comprehension of Text · EMNLP 2016
Ubiquitous computing and smart environments › context recognition
activity recognition
0.212016
Augur: Mining Human Behaviors from Fiction to Power Interactive Systems · CHI 2016
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
commonsense knowledge base
0.112016
Augur: Mining Human Behaviors from Fiction to Power Interactive Systems · CHI 2016
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.112016
SQuAD: 100, 000+ Questions for Machine Comprehension of Text · EMNLP 2016

Methods — techniques the papers use, named apart from their topics

query-code-update · 1.7large language model · 1.7code generation · 1.7multimodal learning · 1.3multi-task learning · 1.3benchmark evaluation · 1.3GPT-4 evaluation · 1.3fine-tuning · 0.9BERT · 0.9backtranslation · 0.4back-translation · 0.4vector model · 0.2text mining · 0.2
YearPublicationVenuePosition
2025 FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models
abstract
Medical vision-language models often struggle with generating accurate quantitative measurements in radiology reports, leading to hallucinations that undermine clinical reliability. We introduce FactCheXcker, a modular framework that de-hallucinates radiology report measurements by leveraging an improved query-code-update paradigm. Specifically, FactCheXcker employs specialized modules and the code generation capabilities of large language models to solve measurement queries generated based on the original report. After extracting measurable findings, the results are incorporated into an updated report. We evaluate FactCheXcker on endotracheal tube placement, which accounts for an average of 78% of report measurements, using the MIMIC-CXR dataset and 11 medical reportgeneration models. Our results show that FactCheXcker significantly reduces hallucinations, improves measurement precision, and maintains the quality of the original reports. Specifically, FactCheXcker improves the performance of all 11 models and achieves an average improvement of 135.0% in reducing measurement hallucinations measured by mean absolute error. Code is available at https://github.com/rajpurkarlab/FactCheXcker.
Alice Heiman, Xiaoman Zhang, Emma Chen, Sung Eun Kim, Pranav Rajpurkar
CVPR5
2023 Exploring the Boundaries of GPT-4 in Radiology
abstract
Qianchu Liu, Stephanie Hyland, Shruthi Bannur, Kenza Bouzid, Daniel Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Khanna, Hoifung Poon, Naoto Usuyama, Anja Thieme, Aditya Nori, Matthew Lungren, Ozan Oktay, Javier Alvarez-Valle. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Qianchu Liu, Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Tajdin Khanna, Hoifung Poon, Naoto Usuyama, Anja Thieme, Aditya V. Nori, Matthew P. Lungren, Ozan Oktay, Javier Alvarez-Valle
EMNLP11
2023 Multimodal Clinical Benchmark for Emergency Care (MC-BEC): A Comprehensive Benchmark for Evaluating Foundation Models in Emergency Medicine
abstract
We propose the Multimodal Clinical Benchmark for Emergency Care (MC-BEC), a comprehensive benchmark for evaluating foundation models in Emergency Medicine using a dataset of 100K+ continuously monitored Emergency Department visits from 2020-2022. MC-BEC focuses on clinically relevant prediction tasks at timescales from minutes to days, including predicting patient decompensation, disposition, and emergency department (ED) revisit, and includes a standardized evaluation framework with train-test splits and evaluation metrics. The multimodal dataset includes a wide range of detailed clinical data, including triage information, prior diagnoses and medications, continuously measured vital signs, electrocardiogram and photoplethysmograph waveforms, orders placed and medications administered throughout the visit, free-text reports of imaging studies, and information on ED diagnosis, disposition, and subsequent revisits. We provide performance baselines for each prediction task to enable the evaluation of multimodal, multitask models. We believe that MC-BEC will encourage researchers to develop more effective, generalizable, and accessible foundation models for multimodal clinical data.
Emma Chen, Aman Kansal, Julie Chen, Boyang Tom Jin, Julia Rachel Reisler, David A. Kim, Pranav Rajpurkar
NeurIPS7
2022 Transfer learning enables prediction of myocardial injury from continuous single-lead electrocardiography
abstract
OBJECTIVE: Chest pain is common, and current risk-stratification methods, requiring 12-lead electrocardiograms (ECGs) and serial biomarker assays, are static and restricted to highly resourced settings. Our objective was to predict myocardial injury using continuous single-lead ECG waveforms similar to those obtained from wearable devices and to evaluate the potential of transfer learning from labeled 12-lead ECGs to improve these predictions. METHODS: We studied 10 874 Emergency Department (ED) patients who received continuous ECG monitoring and troponin testing from 2020 to 2021. We defined myocardial injury as newly elevated troponin in patients with chest pain or shortness of breath. We developed deep learning models of myocardial injury using continuous lead II ECG from bedside monitors as well as conventional 12-lead ECGs from triage. We pretrained single-lead models on a pre-existing corpus of labeled 12-lead ECGs. We compared model predictions to those of ED physicians. RESULTS: A transfer learning strategy, whereby models for continuous single-lead ECGs were first pretrained on 12-lead ECGs from a separate cohort, predicted myocardial injury as accurately as models using patients' own 12-lead ECGs: area under the receiver operating characteristic curve 0.760 (95% confidence interval [CI], 0.721-0.799) and area under the precision-recall curve 0.321 (95% CI, 0.251-0.397). Models demonstrated a high negative predictive value for myocardial injury among patients with chest pain or shortness of breath, exceeding the predictive performance of ED physicians, while attending to known stigmata of myocardial injury. CONCLUSIONS: Deep learning models pretrained on labeled 12-lead ECGs can predict myocardial injury from noisy, continuous monitor data early in a patient's presentation. The utility of continuous single-lead ECG in the risk stratification of chest pain has implications for wearable devices and preclinical settings, where external validation of the approach is needed.
Boyang Tom Jin, Raj Palleti, Siyu Shi, Andrew Y. Ng, James V. Quinn, Pranav Rajpurkar, David A. Kim
J. Am. Medical Informatics Assoc.6
2021 GloFlow: Whole Slide Image Stitching from Video Using Optical Flow and Global Image Alignment
Viswesh Krishna, Anirudh Joshi, Damir Vrabac, Philip L. Bulterys, Eric Yang, Sebastian Fernandez-Pol, Andrew Y. Ng, Pranav Rajpurkar
MICCAI (8)8
2021 Improving hospital readmission prediction using individualized utility analysis
abstract
OBJECTIVE: Machine learning (ML) models for allocating readmission-mitigating interventions are typically selected according to their discriminative ability, which may not necessarily translate into utility in allocation of resources. Our objective was to determine whether ML models for allocating readmission-mitigating interventions have different usefulness based on their overall utility and discriminative ability. MATERIALS AND METHODS: We conducted a retrospective utility analysis of ML models using claims data acquired from the Optum Clinformatics Data Mart, including 513,495 commercially-insured inpatients (mean [SD] age 69 [19] years; 294,895 [57%] Female) over the period January 2016 through January 2017 from all 50 states with mean 90 day cost of $11,552. Utility analysis estimates the cost, in dollars, of allocating interventions for lowering readmission risk based on the reduction in the 90-day cost. RESULTS: Allocating readmission-mitigating interventions based on a GBDT model trained to predict readmissions achieved an estimated utility gain of $104 per patient, and an AUC of 0.76 (95% CI 0.76, 0.77); allocating interventions based on a model trained to predict cost as a proxy achieved a higher utility of $175.94 per patient, and an AUC of 0.62 (95% CI 0.61, 0.62). A hybrid model combining both intervention strategies is comparable with the best models on either metric. Estimated utility varies by intervention cost and efficacy, with each model performing the best under different intervention settings. CONCLUSION: We demonstrate that machine learning models may be ranked differently based on overall utility and discriminative ability. Machine learning models for allocation of limited health resources should consider directly optimizing for utility.
Michael Ko, Emma Chen, Ashwin Agrawal, Pranav Rajpurkar, Anand Avati, Andrew Y. Ng, Sanjay Basu, Nigam H. Shah
J. Biomed. Informatics4
2020 Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT
abstract
The extraction of labels from radiology text reports enables large-scale training of medical imaging models.Existing approaches to report labeling typically rely either on sophisticated feature engineering based on medical domain knowledge or manual annotations by experts.In this work, we introduce a BERTbased approach to medical image report labeling that exploits both the scale of available rule-based systems and the quality of expert annotations.We demonstrate superior performance of a biomedically pretrained BERT model first trained on annotations of a rulebased labeler and then fine-tuned on a small set of expert annotations augmented with automated backtranslation.We find that our final model, CheXbert, is able to outperform the previous best rule-based labeler with statistical significance, setting a new SOTA for report labeling on one of the largest datasets of chest x-rays.
Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, Andrew Y. Ng, Matthew P. Lungren
EMNLP (1)3
2019 CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison
abstract
Large, labeled datasets have driven deep learning methods to achieve expert-level performance on a variety of medical imaging tasks. We present CheXpert, a large dataset that contains 224,316 chest radiographs of 65,240 patients. We design a labeler to automatically detect the presence of 14 observations in radiology reports, capturing uncertainties inherent in radiograph interpretation. We investigate different approaches to using the uncertainty labels for training convolutional neural networks that output the probability of these observations given the available frontal and lateral radiographs. On a validation set of 200 chest radiographic studies which were manually annotated by 3 board-certified radiologists, we find that different uncertainty approaches are useful for different pathologies. We then evaluate our best model on a test set composed of 500 chest radiographic studies annotated by a consensus of 5 board-certified radiologists, and compare the performance of our model to that of 3 additional radiologists in the detection of 5 selected pathologies. On Cardiomegaly, Edema, and Pleural Effusion, the model ROC and PR curves lie above all 3 radiologist operating points. We release the dataset to the public as a standard benchmark to evaluate performance of chest radiograph interpretation models.
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu 0005, Silviana Ciurea-Ilcus, Christopher Chute, Henrik Marklund, Behzad Haghgoo, Robyn L. Ball, Yekaterina Shpanskaya, Jayne Seekins, David A. Mong, Safwan Halabi, Jesse K. Sandberg, Ricky Jones, David B. Larson, Curt Langlotz, Bhavik N. Patel, Matthew P. Lungren, Andrew Y. Ng
AAAI2
2016 Augur: Mining Human Behaviors from Fiction to Power Interactive Systems
abstract
From smart homes that prepare coffee when we wake, to phones that know not to interrupt us during important conversations, our collective visions of HCI imagine a future in which computers understand a broad range of human behaviors. Today our systems fall short of these visions, however, because this range of behaviors is too large for designers or programmers to capture manually. In this paper, we instead demonstrate it is possible to mine a broad knowledge base of human behavior by analyzing more than one billion words of modern fiction. Our resulting knowledge base, Augur, trains vector models that can predict many thousands of user activities from surrounding objects in modern contexts: for example, whether a user may be eating food, meeting with a friend, or taking a selfie. Augur uses these predictions to identify actions that people commonly take on objects in the world and estimate a user's future activities given their current situation. We demonstrate Augur-powered, activity-based systems such as a phone that silences itself when the odds of you answering it are low, and a dynamic music player that adjusts to your present activity. A field deployment of an Augur-powered wearable camera resulted in 96% recall and 71% precision on its unsupervised predictions of common daily activities. A second evaluation where human judges rated the system's predictions over a broad set of input images found that 94% were rated sensible.
Ethan Fast, William McGrath, Pranav Rajpurkar, Michael S. Bernstein
CHI3
2016 SQuAD: 100, 000+ Questions for Machine Comprehension of Text
abstract
We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage. We analyze the dataset to understand the types of reasoning required to answer the questions, leaning heavily on dependency and constituency trees. We build a strong logistic regression model, which achieves an F1 score of 51.0%, a significant improvement over a simple baseline (20%). However, human performance (86.8%) is much higher, indicating that the dataset presents a good challenge problem for future research. The dataset is freely available at https://stanford-qa.com
Pranav Rajpurkar, Jian Zhang 0049, Konstantin Lopyrev, Percy Liang
EMNLP1