EDBT 2026 Demo / reviewers in the wild / expert
Curt Langlotz
dblp:12/1751 · also Curtis Langlotz, Curtis P. Langlotz
· DBLP profile ↗
33ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0002-8972-8051ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated Structured Radiology Report GenerationabstractJean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas, Zhihong Chen, Sophie Ostmeier, Asfandyar Azhar, Kelvin Zhenghao Li, Andrew Johnston, Christian Bluethgen, Eduardo Pontes Reis, Mohamed S Muneer, Maya Varma, Curtis Langlotz. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas, Sophie Ostmeier, Asfandyar Azhar, Kelvin Zhenghao Li, Andrew Johnston, Christian Bluethgen, Eduardo Pontes Reis, Mohamed S. Muneer, Maya Varma, Curt Langlotz |
ACL (1) | 14 |
| 2025 | CheXalign: Preference fine-tuning in chest X-ray interpretation models without human feedbackabstractDennis Hein, Zhihong Chen, Sophie Ostmeier, Justin Xu, Maya Varma, Eduardo Pontes Reis, Arne Edward Michalson Md, Christian Bluethgen, Hyun Joo Shin, Curtis Langlotz, Akshay S Chaudhari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Dennis Hein, Sophie Ostmeier, Justin Xu, Maya Varma, Eduardo Pontes Reis, Arne Edward Michalson, Christian Bluethgen, Hyun Joo Shin, Curt Langlotz, Akshay Chaudhari |
ACL (1) | 10 |
| 2025 | MEDUSA: Medical Data Under Shadow Attacks via Hybrid Model InversionabstractWe introduce MEDUSA (Medical Data Under Shadow Attacks), a novel hybrid model inversion framework that leverages gradient-based optimization and TCNNs to reconstruct high-fidelity medical images from model outputs in a gray-box setting. Unlike traditional attacks requiring full model details, MEDUSA uses surrogate shadow models trained on publicly available data, simulating limited-information scenarios often encountered in practice. Our approach shows that even with restricted access, quality image reconstructions are possible, raising serious privacy concerns for patient data. Contributions include demonstrating that a combination of gradient-based methods and TCNNs yields potent reconstructions, even with limited model access, and providing a detailed analysis of how different input configurations impact reconstruction quality. We also evaluate the reconstructions as viable training data, finding that they can approximate real images well enough to use for model training. Finally, we propose robust defensive mechanisms such as output vector truncation, Gaussian noise, and a new k-NN smearing technique to tackle privacy risks. Asfandyar Azhar, Paul Thielen, Curt Langlotz |
AISTATS | 3 |
| 2025 | Structuring Radiology Reports: Challenging LLMs with Lightweight ModelsabstractJohannes Moll, Louisa Fay, Asfandyar Azhar, Sophie Ostmeier, Sergios Gatidis, Tim C. Lueth, Curtis Langlotz, Jean-Benoit Delbrouck. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Johannes Moll, Louisa Fay, Asfandyar Azhar, Sophie Ostmeier, Sergios Gatidis, Tim C. Lueth, Curt Langlotz, Jean-Benoit Delbrouck |
EMNLP | 7 |
| 2025 | Time-to-Event Pretraining for 3D Medical ImagingabstractWith the rise of medical foundation models and the growing availability of imaging data, scalable pretraining techniques offer a promising way to identify imaging biomarkers predictive of future disease risk. While current self-supervised methods for 3D medical imaging models capture local structural features like organ morphology, they fail to link pixel biomarkers with long-term health outcomes due to a missing context problem. Current approaches lack the temporal context necessary to identify biomarkers correlated with disease progression, as they rely on supervision derived only from images and concurrent text descriptions. To address this, we introduce time-to-event pretraining, a pretraining framework for 3D medical imaging models that leverages large-scale temporal supervision from paired, longitudinal electronic health records (EHRs). Using a dataset of 18,945 CT scans (4.2 million 2D images) and time-to-event distributions across thousands of EHR-derived tasks, our method improves outcome prediction, achieving an average AUROC increase of 23.7% and a 29.4% gain in Harrell’s C-index across 8 benchmark tasks. Importantly, these gains are achieved without sacrificing diagnostic classification performance. This study lays the foundation for integrating longitudinal EHR and 3D imaging data to advance clinical risk prediction. Zepeng Huo, Jason Alan Fries, Alejandro Lozano, Jeya Maria Jose Valanarasu, Ethan Steinberg, Louis Blankemeier, Akshay Chaudhari, Curt Langlotz, Nigam H. Shah |
ICLR | 8 |
| 2025 | LieRE: Lie Rotational Positional EncodingsabstractTransformer architectures rely on position encodings to model the spatial structure of input data. Rotary Position Encoding (RoPE) is a widely used method in language models that encodes relative positions through fixed, block-diagonal, rotation matrices applied to key-query interactions. We hypothesize that this inductive bias limits their RoPE’s effectiveness for modalities with high dimensional structure. Lie Relative Encodings (LieRE) introduce a principled generalization of RoPE, aimed at increasing the representational capacity of positional encodings in transformers. Instead of fixed 2D rotations, LieRE learns dense skew-symmetric matrices (Lie algebra elements), which are then differentiable mapped to form high-dimensional rotation matrices (Lie group elements). This results in richer, learnable, and continuous, encodings of both relative and absolute positional information. We demonstrate the effectiveness of LieRE on 2D and 3D vision tasks, showing that it generalizes well to higher input resolutions while maintaining computational efficiency. The code and checkpoints are publicly available at https://github.com/StanfordMIMI/LieRE. Sophie Ostmeier, Brian Axelrod, Maya Varma, Michael E. Moseley, Akshay Chaudhari, Curt Langlotz |
ICML | 6 |
| 2025 | STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC HistopathologyabstractMulti-class tissue-type classification of colorectal cancer (CRC) histopathologic images is a significant step in the development of downstream machine learning models for diagnosis and treatment planning. However, publicly available CRC datasets used to build tissue classifiers often suffer from insufficient morphologic diversity, class imbalance, and low-quality image tiles, limiting downstream model performance and generalizability. To address this research gap, we introduce STARC-9 (STAnford coloRectal Cancer), a large-scale dataset for multi-class tissue classification. STARC-9 comprises 630,000 histopathologic image tiles uniformly sampled across nine clinically relevant tissue classes (each represented by 70,000 tiles), systematically extracted from hematoxylin & eosin-stained whole-slide images (WSI) from 200 CRC patients at the Stanford University School of Medicine. To construct STARC-9, we propose a novel framework, DeepCluster++, consisting of two primary steps to ensure diversity within each tissue class, followed by pathologist verification. First, an encoder from an autoencoder trained specifically on histopathologic images is used to extract feature vectors from all tiles within a given input WSI. Next, K-means clustering groups morphologically similar tiles, followed by an equal-frequency binning method to sample diverse patterns within each tissue class. Finally, the selected tiles are verified by expert gastrointestinal pathologists to ensure classification accuracy. This semi-automated approach significantly reduces the manual effort required for dataset curation while producing high-quality training examples. To validate the utility of STARC-9, we benchmarked baseline convolutional neural networks, transformers, and pathology-specific foundation models on downstream multi-class CRC tissue classification and segmentation tasks when trained on STARC-9 versus publicly available datasets, demonstrating superior generalizability of models trained on STARC-9. Although we demonstrate the utility of DeepCluster++ on CRC as a pilot use-case, it is a flexible framework that can be used for constructing high-quality datasets from large WSI repositories across a wide range of cancer and non-cancer applications. Barathi Subramanian, Rathinaraja Jeyaraj, Mitchell Nevin Peterson, Terry Guo, Nigam H. Shah, Curt Langlotz, Andrew Y. Ng, Jeanne Shen |
NeurIPS | 6 |
| 2025 | TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language ModelsabstractVision-language models (VLMs) have made great strides in addressing temporal understanding tasks, which involve characterizing visual changes across a sequence of images. However, recent works have suggested that when making predictions, VLMs may rely on static feature biases, such as background or object features, rather than dynamic visual changes. Static feature biases are a type of shortcut and can contribute to systematic prediction errors on downstream tasks; as a result, identifying and characterizing error-inducing static feature biases is critical prior to real-world model deployment. Existing approaches for identifying such systematic failure modes in trained models (i) are typically designed for non-temporal settings and (ii) are challenging to evaluate in temporal settings due to the lack of quantitative evaluation frameworks. In this work, we address these challenges by introducing TRoVe, an automated approach for discovering error-inducing static feature biases learned by temporal VLMs. Given a trained VLM and an annotated validation dataset associated with a downstream classification task, TRoVe extracts candidate static features from the dataset and scores each feature by (i) the effect of the feature on classification errors as well as (ii) the extent to which the VLM relies on the feature when making predictions. In order to quantitatively evaluate TRoVe, we introduce an evaluation framework consisting of 101 trained temporal VLMs paired with ground-truth annotations for learned static feature biases. We use this framework to demonstrate that TRoVe can accurately identify error-inducing static feature biases in VLMs, achieving a 28.6% improvement over the closest baseline. Finally, we apply TRoVe to 7 off-the-shelf VLMs and 2 temporal understanding tasks, surfacing previously-unknown static feature biases and demonstrating that knowledge of learned biases can aid in improving model performance at test time. Our code is available at https://github.com/Stanford-AIMI/TRoVe. Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curt Langlotz |
NeurIPS | 5 |
| 2024 | RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language ModelsabstractFine-tuned vision-language models (VLMs) often capture spurious correlations between image features and textual attributes, resulting in degraded zero-shot performance at test time. Existing approaches for addressing spurious correlations (i) primarily operate at the global image-level rather than intervening directly on fine-grained image features and (ii) are predominantly designed for unimodal settings. In this work, we present RaVL, which takes a fine-grained perspective on VLM robustness by discovering and mitigating spurious correlations using local image features rather than operating at the global image level. Given a fine-tuned VLM, RaVL first discovers spurious correlations by leveraging a region-level clustering approach to identify precise image features contributing to zero-shot classification errors. Then, RaVL mitigates the identified spurious correlation with a novel region-aware loss function that enables the VLM to focus on relevant regions and ignore spurious relationships during fine-tuning. We evaluate RaVL on 654 VLMs with various model architectures, data domains, and learned spurious correlations. Our results show that RaVL accurately discovers (191% improvement over the closest baseline) and mitigates (8.2% improvement on worst-group image classification accuracy) spurious correlations. Qualitative evaluations on general-domain and medical-domain VLMs confirm our findings. Maya Varma, Jean-Benoit Delbrouck, Akshay Chaudhari, Curt Langlotz |
NeurIPS | 5 |
| 2023 | ViLLA: Fine-Grained Vision-Language Representation Learning from Real-World DataabstractVision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are significantly more complex: each image (e.g. X-ray) is often paired with text (e.g. physician report) that describes many distinct attributes occurring in fine-grained regions of the image. We refer to these samples as exhibiting high pairwise complexity, since each image-text pair can be decomposed into a large number of region-attribute pairings. The extent to which VLMs can capture fine-grained relationships between image regions and textual attributes when trained on such data has not been previously evaluated. The first key contribution of this work is to demonstrate through systematic evaluations that as the pairwise complexity of the training dataset increases, standard VLMs struggle to learn region-attribute relationships, exhibiting performance degradations of up to 37% on retrieval tasks. In order to address this issue, we introduce ViLLA as our second key contribution. ViLLA, which is trained to capture fine-grained region-attribute relationships from complex datasets, involves two components: (a) a lightweight, self-supervised mapping model to decompose image-text samples into region-attribute pairs, and (b) a contrastive VLM to learn representations from generated region-attribute pairs. We demonstrate with experiments across four domains (synthetic, product, medical, and natural images) that ViLLA outperforms comparable VLMs on fine-grained reasoning tasks, such as zero-shot object detection (up to 3.6 AP50 points on COCO and 0.6 mAP points on LVIS) and retrieval (up to 14.2 R-Precision points1 Maya Varma, Jean-Benoit Delbrouck, Sarah M. Hooper, Akshay Chaudhari, Curt Langlotz |
ICCV | 5 |
| 2023 | RaLEs: a Benchmark for Radiology Language EvaluationsabstractThe radiology report is the main form of communication between radiologists and other clinicians. Prior work in natural language processing in radiology reports has shown the value of developing methods tailored for individual tasks such as identifying reports with critical results or disease detection. Meanwhile, English and biomedical natural language understanding benchmarks such as the General Language Understanding and Evaluation as well as Biomedical Language Understanding and Reasoning Benchmark have motivated the development of models that can be easily adapted to address many tasks in those domains. Here, we characterize the radiology report as a distinct domain and introduce RaLEs, the Radiology Language Evaluations, as a benchmark for natural language understanding and generation in radiology. RaLEs is comprised of seven natural language understanding and generation evaluations including the extraction of anatomical and disease entities and their relations, procedure selection, and report summarization. We characterize the performance of models designed for the general, biomedical, clinical and radiology domains across these tasks. We find that advances in the general and biomedical domains do not necessarily translate to radiology, and that improved models from the general domain can perform comparably to smaller clinical-specific models. The limited performance of existing pre-trained models on RaLEs highlights the opportunity to improve domain-specific self-supervised models for natural language processing in radiology. We propose RaLEs as a benchmark to promote and track the development of such domain-specific radiology language models. Juanma Zambrano Chaves, Nandita Bhaskhar, Maayane Attias, Jean-Benoit Delbrouck, Daniel L. Rubin, Andreas M. Loening, Curt Langlotz, Akshay Chaudhari |
NeurIPS | 7 |
| 2023 | A case for reframing automated medical image classification as segmentationabstractImage classification and segmentation are common applications of deep learning to radiology. While many tasks can be framed using either classification or segmentation, classification has historically been cheaper to label and more widely used. However, recent work has drastically reduced the cost of training segmentation networks. In light of this recent work, we reexamine the choice of training classification vs. segmentation models. First, we use an information theoretic approach to analyze why segmentation vs. classification models may achieve different performance on the same dataset and overarching task. We then implement multiple methods for using segmentation models to classify medical images, which we call *segmentation-for-classification*, and compare these methods against traditional classification on three retrospective datasets. We use our analysis and experiments to summarize the benefits of switching from segmentation to classification, including: improved sample efficiency, enabling improved performance with fewer labeled images (up to an order of magnitude lower), on low-prevalence classes, and on certain rare subgroups (up to 161.1\% improved recall); improved robustness to spurious correlations (up to 44.8\% improved robust AUROC); and improved model interpretability, evaluation, and error analysis. Sarah M. Hooper, Mayee F. Chen, Khaled Saab 0002, Kush Bhatia, Curt Langlotz, Christopher Ré |
NeurIPS | 5 |
| 2023 | INSPECT: A Multimodal Dataset for Patient Outcome Prediction of Pulmonary EmbolismsabstractSynthesizing information from various data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly available, multimodal medical datasets. To address this limitation, we introduce INSPECT, which contains de-identified longitudinal records from a large cohort of pulmonary embolism (PE) patients, along with ground truth labels for multiple outcomes. INSPECT contains data from 19,402 patients, including CT images, sections of radiology reports, and structured electronic health record (EHR) data (including demographics, diagnoses, procedures, and vitals). Using our provided dataset, we develop and release a benchmark for evaluating several baseline modeling approaches on a variety of important PE related tasks. We evaluate image-only, EHR-only, and fused models. Trained models and the de-identified dataset are made available for non-commercial use under a data use agreement. To the best our knowledge, INSPECT is the largest multimodal dataset for enabling reproducible research on strategies for integrating 3D medical imaging and EHR data. Shih-Cheng Huang, Zepeng Huo, Ethan Steinberg, Chia-Chun Chiang, Curt Langlotz, Matthew P. Lungren, Serena Yeung-Levy, Nigam H. Shah, Jason Alan Fries |
NeurIPS | 5 |
| 2023 | Automated deidentification of radiology reports combining transformer and "hide in plain sight" rule-based methodsabstractOBJECTIVE: To develop an automated deidentification pipeline for radiology reports that detect protected health information (PHI) entities and replaces them with realistic surrogates "hiding in plain sight." MATERIALS AND METHODS: In this retrospective study, 999 chest X-ray and CT reports collected between November 2019 and November 2020 were annotated for PHI at the token level and combined with 3001 X-rays and 2193 medical notes previously labeled, forming a large multi-institutional and cross-domain dataset of 6193 documents. Two radiology test sets, from a known and a new institution, as well as i2b2 2006 and 2014 test sets, served as an evaluation set to estimate model performance and to compare it with previously released deidentification tools. Several PHI detection models were developed based on different training datasets, fine-tuning approaches and data augmentation techniques, and a synthetic PHI generation algorithm. These models were compared using metrics such as precision, recall and F1 score, as well as paired samples Wilcoxon tests. RESULTS: Our best PHI detection model achieves 97.9 F1 score on radiology reports from a known institution, 99.6 from a new institution, 99.5 on i2b2 2006, and 98.9 on i2b2 2014. On reports from a known institution, it achieves 99.1 recall of detecting the core of each PHI span. DISCUSSION: Our model outperforms all deidentifiers it was compared to on all test sets as well as human labelers on i2b2 2014 data. It enables accurate and automatic deidentification of radiology reports. CONCLUSIONS: A transformer-based deidentification pipeline can achieve state-of-the-art performance for deidentifying radiology reports and other medical documents. Pierre J. Chambon, Christopher Wu, Jackson M. Steinkamp, Jason Adleberg, Tessa Sundaram Cook, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Implications of Data As a Public Good: Charting a Path Toward Diagnostic Excellence
Curt Langlotz, Thomas Wang, Tina Hernandez-Boussard, Irene Dankwa-Mullan |
AMIA | 1 |
| 2021 | Cut out the annotator, keep the cutout: better segmentation with weak supervision
Sarah M. Hooper, Michael Wornow, Ying Hang Seah, Peter Kellman, Hui Xue 0006, Frederic Sala, Curt Langlotz, Christopher Ré |
ICLR | 7 |
| 2021 | Improving Factual Completeness and Consistency of Image-to-Text Radiology Report GenerationabstractYasuhide Miura, Yuhao Zhang, Emily Tsai, Curtis Langlotz, Dan Jurafsky. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yasuhide Miura, Yuhao Zhang 0004, Emily Bao Tsai, Curt Langlotz, Daniel Jurafsky |
NAACL-HLT | 4 |
| 2021 | Biomedical and clinical English model packages for the Stanza Python NLP libraryabstractOBJECTIVE: The study sought to develop and evaluate neural natural language processing (NLP) packages for the syntactic analysis and named entity recognition of biomedical and clinical English text. MATERIALS AND METHODS: We implement and train biomedical and clinical English NLP pipelines by extending the widely used Stanza library originally designed for general NLP tasks. Our models are trained with a mix of public datasets such as the CRAFT treebank as well as with a private corpus of radiology reports annotated with 5 radiology-domain entities. The resulting pipelines are fully based on neural networks, and are able to perform tokenization, part-of-speech tagging, lemmatization, dependency parsing, and named entity recognition for both biomedical and clinical text. We compare our systems against popular open-source NLP libraries such as CoreNLP and scispaCy, state-of-the-art models such as the BioBERT models, and winning systems from the BioNLP CRAFT shared task. RESULTS: For syntactic analysis, our systems achieve much better performance compared with the released scispaCy models and CoreNLP models retrained on the same treebanks, and are on par with the winning system from the CRAFT shared task. For NER, our systems substantially outperform scispaCy, and are better or on par with the state-of-the-art performance from BioBERT, while being much more computationally efficient. CONCLUSIONS: We introduce biomedical and clinical NLP packages built for the Stanza library. These packages offer performance that is similar to the state of the art, and are also optimized for ease of use. To facilitate research, we make all our models publicly available. We also provide an online demonstration (http://stanza.run/bio). Yuhao Zhang 0004, Peng Qi 0003, Christopher D. Manning, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology ReportsabstractNeural abstractive summarization models are able to generate summaries which have high overlap with human references.However, existing models are not optimized for factual correctness, a critical metric in real-world applications.In this work, we develop a general framework where we evaluate the factual correctness of a generated summary by factchecking it automatically against its reference using an information extraction module.We further propose a training strategy which optimizes a neural summarization model with a factual correctness reward via reinforcement learning.We apply the proposed method to the summarization of radiology reports, where factual correctness is a key requirement.On two separate datasets collected from hospitals, we show via both automatic and human evaluation that the proposed approach substantially improves the factual correctness and overall quality of outputs over a competitive neural summarization system, producing radiology summaries that approach the quality of humanauthored ones.Background: radiographic examination of the chest.clinical history: 80 years of age, male ... Findings: frontal radiograph of the chest demonstrates repositioning of the right atrial lead possibly into the ivc.... a right apical pneumothorax can be seen from the image.moderate right and small left pleural effusions continue.no pulmonary edema is observed.heart size is upper limits of normal. Human Summary: pneumothorax is seen. bilateral pleural effusions continue.Summary A (ROUGE-L = 0.77): no pneumothorax is observed.bilateral pleural effusions continue.Summary B (ROUGE-L = 0.44): pneumothorax is observed on radiograph.bilateral pleural effusions continue to be seen. Yuhao Zhang 0004, Derek Merck, Emily Bao Tsai, Christopher D. Manning, Curt Langlotz |
ACL | 5 |
| 2019 | CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert ComparisonabstractLarge, labeled datasets have driven deep learning methods to achieve expert-level performance on a variety of medical imaging tasks. We present CheXpert, a large dataset that contains 224,316 chest radiographs of 65,240 patients. We design a labeler to automatically detect the presence of 14 observations in radiology reports, capturing uncertainties inherent in radiograph interpretation. We investigate different approaches to using the uncertainty labels for training convolutional neural networks that output the probability of these observations given the available frontal and lateral radiographs. On a validation set of 200 chest radiographic studies which were manually annotated by 3 board-certified radiologists, we find that different uncertainty approaches are useful for different pathologies. We then evaluate our best model on a test set composed of 500 chest radiographic studies annotated by a consensus of 5 board-certified radiologists, and compare the performance of our model to that of 3 additional radiologists in the detection of 5 selected pathologies. On Cardiomegaly, Edema, and Pleural Effusion, the model ROC and PR curves lie above all 3 radiologist operating points. We release the dataset to the public as a standard benchmark to evaluate performance of chest radiograph interpretation models. Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu 0005, Silviana Ciurea-Ilcus, Christopher Chute, Henrik Marklund, Behzad Haghgoo, Robyn L. Ball, Yekaterina Shpanskaya, Jayne Seekins, David A. Mong, Safwan Halabi, Jesse K. Sandberg, Ricky Jones, David B. Larson, Curt Langlotz, Bhavik N. Patel, Matthew P. Lungren, Andrew Y. Ng |
AAAI | 17 |
| 2019 | Comparative effectiveness of convolutional neural network (CNN) and recurrent neural network (RNN) architectures for radiology text report classification
Imon Banerjee, Yuan Ling, Matthew C. Chen, Sadid A. Hasan, Curt Langlotz, Nathaniel Moradzadeh, Brian E. Chapman, Timothy Amrhein, David A. Mong, Daniel L. Rubin, Oladimeji Farri, Matthew P. Lungren |
Artif. Intell. Medicine | 5 |
| 2019 | Cross-type biomedical named entity recognition with deep multi-task learningabstractMOTIVATION: State-of-the-art biomedical named entity recognition (BioNER) systems often require handcrafted features specific to each entity type, such as genes, chemicals and diseases. Although recent studies explored using neural network models for BioNER to free experts from manual feature engineering, the performance remains limited by the available training data for each entity type. RESULTS: We propose a multi-task learning framework for BioNER to collectively use the training data of different types of entities and improve the performance on each of them. In experiments on 15 benchmark BioNER datasets, our multi-task model achieves substantially better performance compared with state-of-the-art BioNER systems and baseline neural sequence labeling models. Further analysis shows that the large performance gains come from sharing character- and word-level information among relevant biomedical entities across differently labeled corpora. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/yuzhimanhua/lm-lstm-crf. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuan Wang 0008, Yu Zhang 0044, Xiang Ren 0001, Yuhao Zhang 0004, Marinka Zitnik, Jingbo Shang, Curt Langlotz, Jiawei Han 0001 |
Bioinform. | 7 |
| 2018 | Expanding a radiology lexicon using contextual patterns in radiology reportsabstractObjective: Distributional semantics algorithms, which learn vector space representations of words and phrases from large corpora, identify related terms based on contextual usage patterns. We hypothesize that distributional semantics can speed up lexicon expansion in a clinical domain, radiology, by unearthing synonyms from the corpus. Materials and Methods: We apply word2vec, a distributional semantics software package, to the text of radiology notes to identify synonyms for RadLex, a structured lexicon of radiology terms. We stratify performance by term category, term frequency, number of tokens in the term, vector magnitude, and the context window used in vector building. Results: Ranking candidates based on distributional similarity to a target term results in high curation efficiency: on a ranked list of 775 249 terms, >50% of synonyms occurred within the first 25 terms. Synonyms are easier to find if the target term is a phrase rather than a single word, if it occurs at least 100× in the corpus, and if its vector magnitude is between 4 and 5. Some RadLex categories, such as anatomical substances, are easier to identify synonyms for than others. Discussion: The unstructured text of clinical notes contains a wealth of information about human diseases and treatment patterns. However, searching and retrieving information from clinical notes often suffer due to variations in how similar concepts are described in the text. Biomedical lexicons address this challenge, but are expensive to produce and maintain. Distributional semantics algorithms can assist lexicon curation, saving researchers time and money. Bethany Percha, Yuhao Zhang 0004, Selen Bozkurt, Daniel L. Rubin, Russ B. Altman, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | The LOINC RSNA radiology playbook - a unified terminology for radiology proceduresabstractObjective: This paper describes the unified LOINC/RSNA Radiology Playbook and the process by which it was produced. Methods: The Regenstrief Institute and the Radiological Society of North America (RSNA) developed a unification plan consisting of six objectives 1) develop a unified model for radiology procedure names that represents the attributes with an extensible set of values, 2) transform existing LOINC procedure codes into the unified model representation, 3) create a mapping between all the attribute values used in the unified model as coded in LOINC (ie, LOINC Parts) and their equivalent concepts in RadLex, 4) create a mapping between the existing procedure codes in the RadLex Core Playbook and the corresponding codes in LOINC, 5) develop a single integrated governance process for managing the unified terminology, and 6) publicly distribute the terminology artifacts. Results: We developed a unified model and instantiated it in a new LOINC release artifact that contains the LOINC codes and display name (ie LONG_COMMON_NAME) for each procedure, mappings between LOINC and the RSNA Playbook at the procedure code level, and connections between procedure terms and their attribute values that are expressed as LOINC Parts and RadLex IDs. We transformed all the existing LOINC content into the new model and publicly distributed it in standard releases. The organizations have also developed a joint governance process for ongoing maintenance of the terminology. Conclusions: The LOINC/RSNA Radiology Playbook provides a universal terminology standard for radiology orders and results. Daniel J. Vreeman, Swapna Abhyankar, Kenneth C. Wang, Chris Carr, Beverly Collins, Daniel L. Rubin, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 7 |
| 2017 | The LOINC/RSNA Radiology Playbook: A unified terminology for radiology procedures
Daniel J. Vreeman, Ken Wang, Chris Carr, Beverly Collins, Swapna Abhyankar, Jamalynne Deckard, Clement J. McDonald, Daniel L. Rubin, Curt Langlotz |
AMIA | 9 |
| 2016 | Information extraction from multi-institutional radiology reports
Saeed Hassanpour, Curt Langlotz |
Artif. Intell. Medicine | 2 |
| 2016 | Health IT vendors and the academic community: The 2014 ACMI debate
Alexa T. McCray, John P. Glaser, Ross Koppel, Curt Langlotz, Jonathan C. Silverstein |
J. Biomed. Informatics | 4 |
| 2009 | Improving language models for radiology speech recognition
John M. Paulett, Curt Langlotz |
J. Biomed. Informatics | 2 |
| 2000 | Enhancing the expressiveness and usability of structured image reporting systems
Curt Langlotz, L. Meininger |
AMIA | 1 |
| 1986 | A Framework for Comparing Alternative Formalisms for Plausible Reasoning
Eric Horvitz, David Heckerman, Curt Langlotz |
AAAI | 3 |
| 1986 | Using Decision Theory to Justify Heuristics
Curt Langlotz, Edward H. Shortliffe, Lawrence M. Fagan |
AAAI | 1 |
| 1985 | ONYX: An Architecture for Planning in Uncertain Environments
Curt Langlotz, Lawrence M. Fagan, Samson W. Tu, Branimir Sikic |
IJCAI | 1 |
| 1983 | Adapting a Consultation System to Critique User Plans
Curt Langlotz, Edward H. Shortliffe |
Int. J. Man Mach. Stud. | 1 |