EDBT 2026 Demo / reviewers in the wild / expert
Thomas A. Lasko
dblp:88/2487
· DBLP profile ↗
39ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0003-2300-9529ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised discovery of clinical disease signatures using probabilistic independenceabstractOBJECTIVE: This study uses probabilistic independence to disentangle patient-specific sources of disease and their signatures in Electronic Health Record (EHR) data. MATERIALS AND METHODS: We model a disease source as an unobserved root node in the causal graph of observed EHR variables (laboratory test results, medication exposures, billing codes, and demographics), and a signature as the set of downstream effects that a given source has on those observed variables. We used probabilistic independence to infer 2000 sources and their signatures from 9195 variables in 630,000 cross-sectional training instances sampled at random times from 269,099 longitudinal patient records. We evaluated the learned sources by using them to infer and explain the causes of benign vs. malignant pulmonary nodules in 13,252 records, comparing the inferred causes to an external reference list and other medical literature. We compared models trained by three different algorithms and used corresponding models trained directly from the observed variables as baselines. RESULTS: The model recovered 92% of malignant and 30% of benign causes in the reference standard. Of the top 20 inferred causes of malignancy, 14 were not listed in the reference standard, but had supporting evidence in the literature, as did 11 of the top 20 inferred causes of benign nodules. The model decomposed listed malignant causes by an average factor of 5.5 and benign causes by 4.1, with most stratifying by disease course or treatment regimen. Predictive accuracy of causal predictive models trained on source expressions (Random Forest AUC 0.788) was similar to (p = 0.058) their associational baselines (0.738). DISCUSSION: Most of the unrecovered causes were due to the rarity of the condition or lack of sufficient detail in the input data. Surprisingly, the causal model found many patients with apparently undiagnosed cancer as the source of the malignant nodules. Causal model AUC also suggests that some sources remained undiscovered in this cohort. CONCLUSION: These promising results demonstrate the potential of using probabilistic independence to disentangle complex clinical signatures from noisy, asynchronous, and incomplete EHR data that represent the confluence of multiple simultaneous conditions, and to identify patient-specific causes that support precise treatment decisions. Thomas A. Lasko, William W. Stead, John M. Still, Thomas Z. Li, Michael N. Kammer, Marco Barbero Mota, Eric V. Strobl, Bennett A. Landman, Fabien Maldonado |
J. Biomed. Informatics | 1 |
| 2024 | Leveraging explainable artificial intelligence to optimize clinical decision supportabstractOBJECTIVE: To develop and evaluate a data-driven process to generate suggestions for improving alert criteria using explainable artificial intelligence (XAI) approaches. METHODS: We extracted data on alerts generated from January 1, 2019 to December 31, 2020, at Vanderbilt University Medical Center. We developed machine learning models to predict user responses to alerts. We applied XAI techniques to generate global explanations and local explanations. We evaluated the generated suggestions by comparing with alert's historical change logs and stakeholder interviews. Suggestions that either matched (or partially matched) changes already made to the alert or were considered clinically correct were classified as helpful. RESULTS: The final dataset included 2 991 823 firings with 2689 features. Among the 5 machine learning models, the LightGBM model achieved the highest Area under the ROC Curve: 0.919 [0.918, 0.920]. We identified 96 helpful suggestions. A total of 278 807 firings (9.3%) could have been eliminated. Some of the suggestions also revealed workflow and education issues. CONCLUSION: We developed a data-driven process to generate suggestions for improving alert criteria using XAI techniques. Our approach could identify improvements regarding clinical decision support (CDS) that might be overlooked or delayed in manual reviews. It also unveils a secondary purpose for the XAI: to improve quality by discovering scenarios where CDS alerts are not accepted due to workflow, education, or staffing issues. Siru Liu, Allison B. McCoy, Josh F. Peterson, Thomas A. Lasko, Dean F. Sittig, Scott D. Nelson, Jennifer Andrews, Lorraine Patterson, Cheryl M. Cobb, David Mulherin, Colleen T. Morton, Adam Wright |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Longitudinal Multimodal Transformer Integrating Imaging and Latent Clinical Signatures from Routine EHRs for Pulmonary Nodule Classification
Thomas Z. Li, John M. Still, Kaiwen Xu, Ho Hin Lee, Leon Y. Cai, Aravind R. Krishnan, Riqiang Gao, Mirza S. Khan, Sanja Antic, Michael N. Kammer, Kim L. Sandler, Fabien Maldonado, Bennett A. Landman, Thomas A. Lasko |
MICCAI (2) | 14 |
| 2023 | UNesT: Local spatial representation learning with hierarchical transformer for efficient medical segmentation
Xin Yu 0010, Qi Yang 0004, Yinchi Zhou, Leon Y. Cai, Riqiang Gao, Ho Hin Lee, Thomas Z. Li, Shunxing Bao, Zhoubing Xu, Thomas A. Lasko, Richard G. Abramson, Yuankai Huo, Bennett A. Landman, Yucheng Tang |
Medical Image Anal. | 10 |
| 2022 | Validating Data-Driven Clinical Fingerprints as an Input Feature Representation
Marco Barbero Mota, Jorge L. Gamboa, John M. Still, Charles M. Stein, Vivian K. Kawai, Thomas A. Lasko |
AMIA | 6 |
| 2021 | Lung Cancer Risk Estimation with Incomplete Data: A Joint Missing Imputation Perspective
Riqiang Gao, Yucheng Tang, Kaiwen Xu, Ho Hin Lee, Steve Deppen, Kim L. Sandler, Pierre P. Massion, Thomas A. Lasko, Yuankai Huo, Bennett A. Landman |
MICCAI (5) | 8 |
| 2021 | SynTEG: a framework for temporal structured electronic health data simulationabstractOBJECTIVE: Simulating electronic health record data offers an opportunity to resolve the tension between data sharing and patient privacy. Recent techniques based on generative adversarial networks have shown promise but neglect the temporal aspect of healthcare. We introduce a generative framework for simulating the trajectory of patients' diagnoses and measures to evaluate utility and privacy. MATERIALS AND METHODS: The framework simulates date-stamped diagnosis sequences based on a 2-stage process that 1) sequentially extracts temporal patterns from clinical visits and 2) generates synthetic data conditioned on the learned patterns. We designed 3 utility measures to characterize the extent to which the framework maintains feature correlations and temporal patterns in clinical events. We evaluated the framework with billing codes, represented as phenome-wide association study codes (phecodes), from over 500 000 Vanderbilt University Medical Center electronic health records. We further assessed the privacy risks based on membership inference and attribute disclosure attacks. RESULTS: The simulated temporal sequences exhibited similar characteristics to real sequences on the utility measures. Notably, diagnosis prediction models based on real versus synthetic temporal data exhibited an average relative difference in area under the ROC curve of 1.6% with standard deviation of 3.8% for 1276 phecodes. Additionally, the relative difference in the mean occurrence age and time between visits were 4.9% and 4.2%, respectively. The privacy risks in synthetic data, with respect to the membership and attribute inference were negligible. CONCLUSION: This investigation indicates that temporal diagnosis code sequences can be simulated in a manner that provides utility and respects privacy. Ziqi Zhang 0005, Chao Yan 0004, Thomas A. Lasko, Jimeng Sun 0001, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | A Surveillance Framework for Monitoring and Updating Clinical Prediction Models
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 3 |
| 2020 | Detection of calibration drift in clinical prediction models to inform model updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Biomed. Informatics | 3 |
| 2019 | Comparison of Prediction Model Performance Updating Protocols: Using a Data-Driven Testing Procedure to Guide Updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 3 |
| 2019 | Longitudinal modeling of prescription refill records to predict medication use
Kimberley Kondratieff, Candace D. McNaughton, Michael E. Matheny, Thomas A. Lasko |
AMIA | 4 |
| 2019 | Contextual Deep Regression Network for Volume Estimation in Orbital CT
Shikha Chaganti, Camilo Bermudez, Louise A. Mawn, Thomas A. Lasko, Bennett A. Landman |
MICCAI (6) | 4 |
| 2019 | A nonparametric updating method to correct clinical prediction model driftabstractOBJECTIVE: Clinical prediction models require updating as performance deteriorates over time. We developed a testing procedure to select updating methods that minimizes overfitting, incorporates uncertainty associated with updating sample sizes, and is applicable to both parametric and nonparametric models. MATERIALS AND METHODS: We describe a procedure to select an updating method for dichotomous outcome models by balancing simplicity against accuracy. We illustrate the test's properties on simulated scenarios of population shift and 2 models based on Department of Veterans Affairs inpatient admissions. RESULTS: In simulations, the test generally recommended no update under no population shift, no update or modest recalibration under case mix shifts, intercept correction under changing outcome rates, and refitting under shifted predictor-outcome associations. The recommended updates provided superior or similar calibration to that achieved with more complex updating. In the case study, however, small update sets lead the test to recommend simpler updates than may have been ideal based on subsequent performance. DISCUSSION: Our test's recommendations highlighted the benefits of simple updating as opposed to systematic refitting in response to performance drift. The complexity of recommended updating methods reflected sample size and magnitude of performance drift, as anticipated. The case study highlights the conservative nature of our test. CONCLUSIONS: This new test supports data-driven updating of models developed with both biostatistical and machine learning approaches, promoting the transportability and maintenance of a wide array of clinical prediction models and, in turn, a variety of applications relying on modern prediction tools. Sharon E. Davis, Robert A. Greevy Jr., Christopher Fonnesbeck, Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Cost-aware active learning for named entity recognition in clinical textabstractOBJECTIVE: Active Learning (AL) attempts to reduce annotation cost (ie, time) by selecting the most informative examples for annotation. Most approaches tacitly (and unrealistically) assume that the cost for annotating each sample is identical. This study introduces a cost-aware AL method, which simultaneously models both the annotation cost and the informativeness of the samples and evaluates both via simulation and user studies. MATERIALS AND METHODS: We designed a novel, cost-aware AL algorithm (Cost-CAUSE) for annotating clinical named entities; we first utilized lexical and syntactic features to estimate annotation cost, then we incorporated this cost measure into an existing AL algorithm. Using the 2010 i2b2/VA data set, we then conducted a simulation study comparing Cost-CAUSE with noncost-aware AL methods, and a user study comparing Cost-CAUSE with passive learning. RESULTS: Our cost model fit empirical annotation data well, and Cost-CAUSE increased the simulation area under the learning curve (ALC) scores by up to 5.6% and 4.9%, compared with random sampling and alternate AL methods. Moreover, in a user annotation task, Cost-CAUSE outperformed passive learning on the ALC score and reduced annotation time by 20.5%-30.2%. DISCUSSION: Although AL has proven effective in simulations, our user study shows that a real-world environment is far more complex. Other factors have a noticeable effect on the AL method, such as the annotation accuracy of users, the tiredness of users, and even the physical and mental condition of users. CONCLUSION: Cost-CAUSE saves significant annotation cost compared to random sampling. Qiang Wei 0002, Yukun Chen 0001, Mandana Salimi, Joshua C. Denny, Qiaozhu Mei, Thomas A. Lasko, Qingxia Chen, Stephen Wu 0004, Amy Franklin, Trevor Cohen, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Electronic Medical Record Context Signatures Improve Diagnostic Classification Using Medical Image ComputingabstractComposite models that combine medical imaging with electronic medical records (EMR) improve predictive power when compared to traditional models that use imaging alone. The digitization of EMR provides potential access to a wealth of medical information, but presents new challenges in algorithm design and inference. Previous studies, such as Phenome Wide Association Study (PheWAS), have shown that EMR data can be used to investigate the relationship between genotypes and clinical conditions. Here, we introduce Phenome-Disease Association Study to extend the statistical capabilities of the PheWAS software through a custom Python package, which creates diagnostic EMR signatures to capture system-wide co-morbidities for a disease population within a given time interval. We investigate the effect of integrating these EMR signatures with radiological data to improve diagnostic classification in disease domains known to have confounding factors because of variable and complex clinical presentation. Specifically, we focus on two studies: First, a study of four major optic nerve related conditions; and second, a study of diabetes. Addition of EMR signature vectors to radiologically derived structural metrics improves the area under the curve (AUC) for diagnostic classification using elastic net regression, for diseases of the optic nerve. For glaucoma, the AUC improves from 0.71 to 0.83, for intrinsic optic nerve disease it increases from 0.72 to 0.91, for optic nerve edema it increases from 0.95 to 0.96, and for thyroid eye disease from 0.79 to 0.89. The EMR signatures recapitulate known comorbidities with diabetes, such as abnormal glucose, but do not significantly modulate image-derived features. In summary, EMR signatures present a scalable and readily applicable. Shikha Chaganti, Louise A. Mawn, Hakmook Kang, Josephine Egan, Susan M. Resnick, Lori L. Beason-Held, Bennett A. Landman, Thomas A. Lasko |
IEEE J. Biomed. Health Informatics | 8 |
| 2018 | Crowdsourcing Clinical Chart Reviews
Joseph R. Coco, Cheng Ye 0001, Chen Hajaj, Yevgeniy Vorobeychik, Joshua C. Denny, Laurie L. Novak, Bradley A. Malin, Thomas A. Lasko, Daniel Fabbri |
AMIA | 8 |
| 2018 | Automated Mapping of Laboratory Tests to LOINC Codes using Noisy Labels in a National Electronic Health Record System Database
Sharidan K. Parr, Thomas A. Lasko, Alvin D. Jeffery, Matthew S. Shotwell, Michael E. Matheny |
AMIA | 2 |
| 2018 | Automated mapping of laboratory tests to LOINC codes using noisy labels in a national electronic health record system databaseabstractObjective: Standards such as the Logical Observation Identifiers Names and Codes (LOINC®) are critical for interoperability and integrating data into common data models, but are inconsistently used. Without consistent mapping to standards, clinical data cannot be harmonized, shared, or interpreted in a meaningful context. We sought to develop an automated machine learning pipeline that leverages noisy labels to map laboratory data to LOINC codes. Materials and Methods: Across 130 sites in the Department of Veterans Affairs Corporate Data Warehouse, we selected the 150 most commonly used laboratory tests with numeric results per site from 2000 through 2016. Using source data text and numeric fields, we developed a machine learning model and manually validated random samples from both labeled and unlabeled datasets. Results: The raw laboratory data consisted of >6.5 billion test results, with 2215 distinct LOINC codes. The model predicted the correct LOINC code in 85% of the unlabeled data and 96% of the labeled data by test frequency. In the subset of labeled data where the original and model-predicted LOINC codes disagreed, the model-predicted LOINC code was correct in 83% of the data by test frequency. Conclusion: Using a completely automated process, we are able to assign LOINC codes to unlabeled data with high accuracy. When the model-predicted LOINC code differed from the original LOINC code, the model prediction was correct in the vast majority of cases. This scalable, automated algorithm may improve data quality and interoperability, while substantially reducing the manual effort currently needed to accurately map laboratory data. Sharidan K. Parr, Matthew S. Shotwell, Alvin D. Jeffery, Thomas A. Lasko, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 4 |
| 2017 | Calibration Drift Among Regression and Machine Learning Models for Hospital Mortality
Sharon E. Davis, Thomas A. Lasko, Guanhua Chen 0002, Michael E. Matheny |
AMIA | 2 |
| 2017 | Predicting Medications from Diagnostic Codes with Recurrent Neural Networks
Jacek M. Bajor, Thomas A. Lasko |
ICLR (Poster) | 2 |
| 2017 | Calibration drift in regression and machine learning models for acute kidney injuryabstractOBJECTIVE: Predictive analytics create opportunities to incorporate personalized risk estimates into clinical decision support. Models must be well calibrated to support decision-making, yet calibration deteriorates over time. This study explored the influence of modeling methods on performance drift and connected observed drift with data shifts in the patient population. MATERIALS AND METHODS: Using 2003 admissions to Department of Veterans Affairs hospitals nationwide, we developed 7 parallel models for hospital-acquired acute kidney injury using common regression and machine learning methods, validating each over 9 subsequent years. RESULTS: Discrimination was maintained for all models. Calibration declined as all models increasingly overpredicted risk. However, the random forest and neural network models maintained calibration across ranges of probability, capturing more admissions than did the regression models. The magnitude of overprediction increased over time for the regression models while remaining stable and small for the machine learning models. Changes in the rate of acute kidney injury were strongly linked to increasing overprediction, while changes in predictor-outcome associations corresponded with diverging patterns of calibration drift across methods. CONCLUSIONS: Efficient and effective updating protocols will be essential for maintaining accuracy of, user confidence in, and safety of personalized risk predictions to support decision-making. Model updating protocols should be tailored to account for variations in calibration drift across methods and respond to periods of rapid performance drift rather than be limited to regularly scheduled annual or biannual intervals. Sharon E. Davis, Thomas A. Lasko, Guanhua Chen 0002, Edward D. Siew, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Evaluating electronic health record data sources and algorithmic approaches to identify hypertensive individualsabstractOBJECTIVE: Phenotyping algorithms applied to electronic health record (EHR) data enable investigators to identify large cohorts for clinical and genomic research. Algorithm development is often iterative, depends on fallible investigator intuition, and is time- and labor-intensive. We developed and evaluated 4 types of phenotyping algorithms and categories of EHR information to identify hypertensive individuals and controls and provide a portable module for implementation at other sites. MATERIALS AND METHODS: We reviewed the EHRs of 631 individuals followed at Vanderbilt for hypertension status. We developed features and phenotyping algorithms of increasing complexity. Input categories included International Classification of Diseases, Ninth Revision (ICD9) codes, medications, vital signs, narrative-text search results, and Unified Medical Language System (UMLS) concepts extracted using natural language processing (NLP). We developed a module and tested portability by replicating 10 of the best-performing algorithms at the Marshfield Clinic. RESULTS: Random forests using billing codes, medications, vitals, and concepts had the best performance with a median area under the receiver operator characteristic curve (AUC) of 0.976. Normalized sums of all 4 categories also performed well (0.959 AUC). The best non-NLP algorithm combined normalized ICD9 codes, medications, and blood pressure readings with a median AUC of 0.948. Blood pressure cutoffs or ICD9 code counts alone had AUCs of 0.854 and 0.908, respectively. Marshfield Clinic results were similar. CONCLUSION: This work shows that billing codes or blood pressure readings alone yield good hypertension classification performance. However, even simple combinations of input categories improve performance. The most complex algorithms classified hypertension with excellent recall and precision. Pedro L. Teixeira, Wei-Qi Wei, Robert M. Cronin, Huan Mo, Jacob P. VanHouten, Robert J. Carroll, Eric LaRose, Lisa Bastarache, S. Trent Rosenbloom, Todd L. Edwards, Dan M. Roden, Thomas A. Lasko, Richard A. Dart, Anne M. Nikolai, Peggy L. Peissig, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 12 |
| 2016 | The Discriminative Power of Non-Specific Laboratory Results
Jacob P. VanHouten, Christopher Fonnesbeck, Michael E. Matheny, Thomas A. Lasko |
AMIA | 4 |
| 2015 | Nonstationary Gaussian Process Regression for Evaluating Repeated Clinical Laboratory TestsabstractSampling repeated clinical laboratory tests with appropriate timing is challenging because the latent physiologic function being sampled is in general nonstationary. When ordering repeated tests, clinicians adopt various simple strategies that may or may not be well suited to the behavior of the function. Previous research on this topic has been primarily focused on cost-driven assessments of oversampling. But for monitoring physiologic state or for retrospective analysis, undersampling can be much more problematic than oversampling. In this paper we analyze hundreds of observation sequences of four different clinical laboratory tests to provide principled, data-driven estimates of undersampling and oversampling, and to assess whether the sampling adapts to changing volatility of the latent function. To do this, we developed a new method for fitting a Gaussian process to samples of a nonstationary latent function. Our method includes an explicit estimate of the latent function's volatility over time, which is deterministically related to its nonstationarity. We find on average that the degree of undersampling is up to an order of magnitude greater than oversampling, and that only a small minority are sampled with an adaptive strategy. Thomas A. Lasko |
AAAI | 1 |
| 2015 | Assessing Variability in Breast Cancer Treatment Paths Using Frequent Sequence Mining
Ravi V. Atreya, Thomas A. Lasko, Mia A. Levy |
AMIA | 2 |
| 2015 | Real Time Active Learning Study for Clinical Named Entity Recognition
Yukun Chen 0001, Sungrim Moon, Thomas A. Lasko, Qiaozhu Mei, Trevor Cohen, Qingxia Chen, Joshua C. Denny, Hua Xu 0001 |
AMIA | 3 |
| 2015 | A study of active learning methods for named entity recognition in clinical text
Yukun Chen 0001, Thomas A. Lasko, Qiaozhu Mei, Joshua C. Denny, Hua Xu 0001 |
J. Biomed. Informatics | 2 |
| 2014 | Machine Learning for Risk Prediction of Acute Coronary Syndrome
Jacob P. VanHouten, Jack Starmer, Nancy M. Lorenzi, David J. Maron, Thomas A. Lasko |
AMIA | 5 |
| 2014 | Efficient Inference of Gaussian-Process-Modulated Renewal Processes with Application to Medical Event Data
Thomas A. Lasko |
UAI | 1 |
| 2014 | Predicting changes in hypertension control using electronic health records from a chronic disease management programabstractOBJECTIVE: Common chronic diseases such as hypertension are costly and difficult to manage. Our ultimate goal is to use data from electronic health records to predict the risk and timing of deterioration in hypertension control. Towards this goal, this work predicts the transition points at which hypertension is brought into, as well as pushed out of, control. METHOD: In a cohort of 1294 patients with hypertension enrolled in a chronic disease management program at the Vanderbilt University Medical Center, patients are modeled as an array of features derived from the clinical domain over time, which are distilled into a core set using an information gain criteria regarding their predictive performance. A model for transition point prediction was then computed using a random forest classifier. RESULTS: The most predictive features for transitions in hypertension control status included hypertension assessment patterns, comorbid diagnoses, procedures and medication history. The final random forest model achieved a c-statistic of 0.836 (95% CI 0.830 to 0.842) and an accuracy of 0.773 (95% CI 0.766 to 0.780). CONCLUSIONS: This study achieved accurate prediction of transition points of hypertension control status, an important first step in the long-term goal of developing personalized hypertension management plans. Jimeng Sun 0001, Candace D. McNaughton, Ping Zhang 0016, Adam Perer, Aris Gkoulalas-Divanis, Joshua C. Denny, Jacqueline Kirby, Thomas A. Lasko, Alexander Saip, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 8 |
| 2013 | Role of ICD Granularity in Phenotyping Hematologic Malignancies for Tumor Registries
Ravi V. Atreya, Thomas A. Lasko, Mia A. Levy |
AMIA | 2 |
| 2013 | A Study of Active Learning Methods for Clinical Entities Recognition
Yukun Chen 0001, Thomas A. Lasko, Qiaozhu Mei, Joshua C. Denny, Hua Xu 0001 |
AMIA | 2 |
| 2013 | Random Forest Classification of Acute Coronary Syndromes
Jacob P. VanHouten, Jack Starmer, Nancy M. Lorenzi, Thomas A. Lasko |
AMIA | 4 |
| 2013 | Development and evaluation of an ensemble resource linking medications to their indicationsabstractOBJECTIVE: To create a computable MEDication Indication resource (MEDI) to support primary and secondary use of electronic medical records (EMRs). MATERIALS AND METHODS: We processed four public medication resources, RxNorm, Side Effect Resource (SIDER) 2, MedlinePlus, and Wikipedia, to create MEDI. We applied natural language processing and ontology relationships to extract indications for prescribable, single-ingredient medication concepts and all ingredient concepts as defined by RxNorm. Indications were coded as Unified Medical Language System (UMLS) concepts and International Classification of Diseases, 9th edition (ICD9) codes. A total of 689 extracted indications were randomly selected for manual review for accuracy using dual-physician review. We identified a subset of medication-indication pairs that optimizes recall while maintaining high precision. RESULTS: MEDI contains 3112 medications and 63 343 medication-indication pairs. Wikipedia was the largest resource, with 2608 medications and 34 911 pairs. For each resource, estimated precision and recall, respectively, were 94% and 20% for RxNorm, 75% and 33% for MedlinePlus, 67% and 31% for SIDER 2, and 56% and 51% for Wikipedia. The MEDI high-precision subset (MEDI-HPS) includes indications found within either RxNorm or at least two of the three other resources. MEDI-HPS contains 13 304 unique indication pairs regarding 2136 medications. The mean±SD number of indications for each medication in MEDI-HPS is 6.22 ± 6.09. The estimated precision of MEDI-HPS is 92%. CONCLUSIONS: MEDI is a publicly available, computable resource that links medications with their indications as represented by concepts and billing codes. MEDI may benefit clinical EMR applications and reuse of EMR data for research. Wei-Qi Wei, Robert M. Cronin, Hua Xu 0001, Thomas A. Lasko, Lisa Bastarache, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | Portability of an algorithm to identify rheumatoid arthritis in electronic health recordsabstractOBJECTIVES: Electronic health records (EHR) can allow for the generation of large cohorts of individuals with given diseases for clinical and genomic research. A rate-limiting step is the development of electronic phenotype selection algorithms to find such cohorts. This study evaluated the portability of a published phenotype algorithm to identify rheumatoid arthritis (RA) patients from EHR records at three institutions with different EHR systems. MATERIALS AND METHODS: Physicians reviewed charts from three institutions to identify patients with RA. Each institution compiled attributes from various sources in the EHR, including codified data and clinical narratives, which were searched using one of two natural language processing (NLP) systems. The performance of the published model was compared with locally retrained models. RESULTS: Applying the previously published model from Partners Healthcare to datasets from Northwestern and Vanderbilt Universities, the area under the receiver operating characteristic curve was found to be 92% for Northwestern and 95% for Vanderbilt, compared with 97% at Partners. Retraining the model improved the average sensitivity at a specificity of 97% to 72% from the original 65%. Both the original logistic regression models and locally retrained models were superior to simple billing code count thresholds. DISCUSSION: These results show that a previously published algorithm for RA is portable to two external hospitals using different EHR systems, different NLP systems, and different target NLP vocabularies. Retraining the algorithm primarily increased the sensitivity at each site. CONCLUSION: Electronic phenotype algorithms allow rapid identification of case populations in multiple sites with little retraining. Robert J. Carroll, William K. Thompson, Anne E. Eyler, Arthur M. Mandelin, Tianxi Cai, Raquel M. Zink, Jennifer A. Pacheco, Chad S. Boomershine, Thomas A. Lasko, Hua Xu 0001, Elizabeth W. Karlson, Raúl G. Pérez, Vivian S. Gainer, Shawn N. Murphy, Eric M. Ruderman, Richard M. Pope, Robert M. Plenge, Abel N. Kho, Katherine P. Liao, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 9 |
| 2010 | Spectral Anonymization of DataabstractThe goal of data anonymization is to allow the release of scientifically useful data in a form that protects the privacy of its subjects. This requires more than simply removing personal identifiers from the data, because an attacker can still use auxiliary information to infer sensitive individual information. Additional perturbation is necessary to prevent these inferences, and the challenge is to perturb the data in a way that preserves its analytic utility.No existing anonymization algorithm provides both perfect privacy protection and perfect analytic utility. We make the new observation that anonymization algorithms are not required to operate in the original vector-space basis of the data, and many algorithms can be improved by operating in a judiciously chosen alternate basis. A spectral basis derived from the data's eigenvectors is one that can provide substantial improvement. We introduce the term spectral anonymization to refer to an algorithm that uses a spectral basis for anonymization, and we give two illustrative examples.We also propose new measures of privacy protection that are more general and more informative than existing measures, and a principled reference standard with which to define adequate privacy protection. Thomas A. Lasko, Staal Amund Vinterbo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2006 | Research Paper: Automated Identification of a Physician's Primary PatientsabstractOBJECTIVE: To develop and validate an automated method for determining the set of patients for whom a given primary care physician holds overall clinical responsibility. DESIGN: The study included all adult patients (16,185) seen at least once in an ambulatory setting during a three-year period by 18 primary care physicians in ten practices. The physicians indicated whether they considered themselves to be the physician primarily responsible for the overall clinical care of each visiting patient. Statistical models were constructed to predict the physicians' designations using predictor variables derived from electronically available appointment schedules and demographic information. MEASUREMENTS: Predictive accuracy was assessed primarily using the area under the receiver-operating characteristic curve (AUC), and secondarily using positive predictive value (PPV) and sensitivity. RESULTS: A minimal set of six variables was identified as predictive of the physicians' designations. The constructed model had a median AUC for individual physicians of 0.92 (interquartile interval: 0.90-0.96), a PPV of 0.94 (interquartile interval: 0.87-0.95), and a sensitivity of 0.95 (interquartile interval: 0.87-0.97). CONCLUSION: A statistical model using a minimal set of commonly available electronic data can accurately predict the set of patients for whom a physician holds primary clinical responsibility. Further research examining the generalization of the model to other settings would be valuable. Thomas A. Lasko, Steven J. Atlas, Michael J. Barry, Henry C. Chueh |
J. Am. Medical Informatics Assoc. | 1 |
| 2005 | The use of receiver operating characteristic curves in biomedical informatics
Thomas A. Lasko, Jui G. Bhagwat, Kelly H. Zou, Lucila Ohno-Machado |
J. Biomed. Informatics | 1 |
| 2002 | DXplain Evoking Strength - Clinician Interpretation and Consistency
Thomas A. Lasko, Mitchell J. Feldman, G. Octo Barnett |
AMIA | 1 |