VLDB 2026 Research / reviewers in the wild / expert
Adler J. Perotte
dblp:66/2491
· DBLP profile ↗
38ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-6695-0282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accurate and Scalable Stochastic Gaussian Process Regression via Learnable Coreset-based Variational InferenceabstractWe introduce a novel stochastic variational inference method for Gaussian process ($\mathcal{GP}$) regression, by deriving a posterior over a learnable set of coresets: i.e., over pseudo-input/output, weighted pairs. Unlike former free-form variational families for stochastic inference, our coreset-based variational $\mathcal{GP}$ (CVGP) is defined in terms of the $\mathcal{GP}$ prior and the (weighted) data likelihood. This formulation naturally incorporates inductive biases of the prior, and ensures its kernel and likelihood dependencies are shared with the posterior. We derive a variational lower-bound on the log-marginal likelihood by marginalizing over the latent $\mathcal{GP}$ coreset variables, and show that CVGP’s lower-bound is amenable to stochastic optimization. CVGP reduces the dimensionality of the variational parameter search space to linear $\mathcal{O}(M)$ complexity, while ensuring numerical stability at $\mathcal{O}(M^3)$ time complexity and $\mathcal{O}(M^2)$ space complexity. Evaluations on real-world and simulated regression problems demonstrate that CVGP achieves superior inference and predictive performance than state-of-the-art, stochastic sparse $\mathcal{GP}$ approximation methods. Mert Ketenci, Adler J. Perotte, Noémie Elhadad, Iñigo Urteaga |
UAI | 2 |
| 2023 | EvidenceMap: a three-level knowledge representation for medical evidence computation and comprehensionabstractOBJECTIVE: To develop a computable representation for medical evidence and to contribute a gold standard dataset of annotated randomized controlled trial (RCT) abstracts, along with a natural language processing (NLP) pipeline for transforming free-text RCT evidence in PubMed into the structured representation. MATERIALS AND METHODS: Our representation, EvidenceMap, consists of 3 levels of abstraction: Medical Evidence Entity, Proposition and Map, to represent the hierarchical structure of medical evidence composition. Randomly selected RCT abstracts were annotated following EvidenceMap based on the consensus of 2 independent annotators to train an NLP pipeline. Via a user study, we measured how the EvidenceMap improved evidence comprehension and analyzed its representative capacity by comparing the evidence annotation with EvidenceMap representation and without following any specific guidelines. RESULTS: Two corpora including 229 disease-agnostic and 80 COVID-19 RCT abstracts were annotated, yielding 12 725 entities and 1602 propositions. EvidenceMap saves users 51.9% of the time compared to reading raw-text abstracts. Most evidence elements identified during the freeform annotation were successfully represented by EvidenceMap, and users gave the enrollment, study design, and study Results sections mean 5-scale Likert ratings of 4.85, 4.70, and 4.20, respectively. The end-to-end evaluations of the pipeline show that the evidence proposition formulation achieves F1 scores of 0.84 and 0.86 in the adjusted random index score. CONCLUSIONS: EvidenceMap extends the participant, intervention, comparator, and outcome framework into 3 levels of abstraction for transforming free-text evidence from the clinical literature into a computable structure. It can be used as an interoperable format for better evidence retrieval and synthesis and an interpretable representation to efficiently comprehend RCT findings. Tian Kang, Yingcheng Sun, Jae Hyun Kim, Casey N. Ta, Adler J. Perotte, Kayla Schiffer, Mutong Wu, Nour Fahmy, Yifan Peng 0002, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | LVHNet: Detecting Cardiac Structural Abnormalities with Chest X-Rays
Shreyas Bhave, Pierre A. Elias, Victor Alfonso Rodriguez, Tim Poterucha, Simi Mutasa, Jay Leb, Nir Uriel, Adler J. Perotte |
AMIA | 8 |
| 2021 | Multi-Channel LSTM for Modeling Irregularly Sampled Time Series
Shreyas Bhave, Adler J. Perotte |
AMIA | 2 |
| 2021 | Inverse-Weighted Survival GamesabstractDeep models trained through maximum likelihood have achieved state-of-the-art results for survival analysis. Despite this training scheme, practitioners evaluate models under other criteria, such as binary classification losses at a chosen set of time horizons, e.g. Brier score (BS) and Bernoulli log likelihood (BLL). Models trained with maximum likelihood may have poor BS or BLL since maximum likelihood does not directly optimize these criteria. Directly optimizing criteria like BS requires inverse-weighting by the censoring distribution. However, estimating the censoring model under these metrics requires inverse-weighting by the failure distribution. The objective for each model requires the other, but neither are known. To resolve this dilemma, we introduce Inverse-Weighted Survival Games. In these games, objectives for each model are built from re-weighted estimates featuring the other model, where the latter is held fixed during training. When the loss is proper, we show that the games always have the true failure and censoring distributions as a stationary point. This means models in the game do not leave the correct distributions once reached. We construct one case where this stationary point is unique. We show that these games optimize BS on simulations and then apply these principles on real world cancer and critically-ill patient data. Xintian Han, Mark Goldstein, Aahlad Manas Puli, Thomas Wies, Adler J. Perotte, Rajesh Ranganath |
NeurIPS | 5 |
| 2021 | Utilizing timestamps of longitudinal electronic health record data to classify clinical deterioration eventsabstractOBJECTIVE: To propose an algorithm that utilizes only timestamps of longitudinal electronic health record data to classify clinical deterioration events. MATERIALS AND METHODS: This retrospective study explores the efficacy of machine learning algorithms in classifying clinical deterioration events among patients in intensive care units using sequences of timestamps of vital sign measurements, flowsheets comments, order entries, and nursing notes. We design a data pipeline to partition events into discrete, regular time bins that we refer to as timesteps. Logistic regressions, random forest classifiers, and recurrent neural networks are trained on datasets of different length of timesteps, respectively, against a composite outcome of death, cardiac arrest, and Rapid Response Team calls. Then these models are validated on a holdout dataset. RESULTS: A total of 6720 intensive care unit encounters meet the criteria and the final dataset includes 830 578 timestamps. The gated recurrent unit model utilizes timestamps of vital signs, order entries, flowsheet comments, and nursing notes to achieve the best performance on the time-to-outcome dataset, with an area under the precision-recall curve of 0.101 (0.06, 0.137), a sensitivity of 0.443, and a positive predictive value of 0. 092 at the threshold of 0.6. DISCUSSION AND CONCLUSION: This study demonstrates that our recurrent neural network models using only timestamps of longitudinal electronic health record data that reflect healthcare processes achieve well-performing discriminative power. Li-heng Fu, Christopher Knaplund, Kenrick Cato, Adler J. Perotte, Min-Jeoung Kang, Patricia C. Dykes, David J. Albers, Sarah Collins Rossetti |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | UMLS-based data augmentation for natural language processing of clinical research literatureabstractOBJECTIVE: The study sought to develop and evaluate a knowledge-based data augmentation method to improve the performance of deep learning models for biomedical natural language processing by overcoming training data scarcity. MATERIALS AND METHODS: We extended the easy data augmentation (EDA) method for biomedical named entity recognition (NER) by incorporating the Unified Medical Language System (UMLS) knowledge and called this method UMLS-EDA. We designed experiments to systematically evaluate the effect of UMLS-EDA on popular deep learning architectures for both NER and classification. We also compared UMLS-EDA to BERT. RESULTS: UMLS-EDA enables substantial improvement for NER tasks from the original long short-term memory conditional random fields (LSTM-CRF) model (micro-F1 score: +5%, + 17%, and +15%), helps the LSTM-CRF model (micro-F1 score: 0.66) outperform LSTM-CRF with transfer learning by BERT (0.63), and improves the performance of the state-of-the-art sentence classification model. The largest gain on micro-F1 score is 9%, from 0.75 to 0.84, better than classifiers with BERT pretraining (0.82). CONCLUSIONS: This study presents a UMLS-based data augmentation method, UMLS-EDA. It is effective at improving deep learning models for both NER and sentence classification, and contributes original insights for designing new, superior deep learning approaches for low-resource biomedical domains. Tian Kang, Adler J. Perotte, Youlan Tang, Casey N. Ta, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | A neuro-symbolic method for understanding free-text medical evidenceabstractOBJECTIVE: We introduce Medical evidence Dependency (MD)-informed attention, a novel neuro-symbolic model for understanding free-text clinical trial publications with generalizability and interpretability. MATERIALS AND METHODS: We trained one head in the multi-head self-attention model to attend to the Medical evidence Ddependency (MD) and to pass linguistic and domain knowledge on to later layers (MD informed). This MD-informed attention model was integrated into BioBERT and tested on 2 public machine reading comprehension benchmarks for clinical trial publications: Evidence Inference 2.0 and PubMedQA. We also curated a small set of recently published articles reporting randomized controlled trials on COVID-19 (coronavirus disease 2019) following the Evidence Inference 2.0 guidelines to evaluate the model's robustness to unseen data. RESULTS: The integration of MD-informed attention head improves BioBERT substantially in both benchmark tasks-as large as an increase of +30% in the F1 score-and achieves the new state-of-the-art performance on the Evidence Inference 2.0. It achieves 84% and 82% in overall accuracy and F1 score, respectively, on the unseen COVID-19 data. CONCLUSIONS: MD-informed attention empowers neural reading comprehension models with interpretability and generalizability via reusable domain knowledge. Its compositionality can benefit any transformer-based architecture for machine reading comprehension of free-text medical evidence. Tian Kang, Ali Turfah, Adler J. Perotte, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Development and validation of prediction models for mechanical ventilation, renal replacement therapy, and readmission in COVID-19 patientsabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) patients are at risk for resource-intensive outcomes including mechanical ventilation (MV), renal replacement therapy (RRT), and readmission. Accurate outcome prognostication could facilitate hospital resource allocation. We develop and validate predictive models for each outcome using retrospective electronic health record data for COVID-19 patients treated between March 2 and May 6, 2020. MATERIALS AND METHODS: For each outcome, we trained 3 classes of prediction models using clinical data for a cohort of SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2)-positive patients (n = 2256). Cross-validation was used to select the best-performing models per the areas under the receiver-operating characteristic and precision-recall curves. Models were validated using a held-out cohort (n = 855). We measured each model's calibration and evaluated feature importances to interpret model output. RESULTS: The predictive performance for our selected models on the held-out cohort was as follows: area under the receiver-operating characteristic curve-MV 0.743 (95% CI, 0.682-0.812), RRT 0.847 (95% CI, 0.772-0.936), readmission 0.871 (95% CI, 0.830-0.917); area under the precision-recall curve-MV 0.137 (95% CI, 0.047-0.175), RRT 0.325 (95% CI, 0.117-0.497), readmission 0.504 (95% CI, 0.388-0.604). Predictions were well calibrated, and the most important features within each model were consistent with clinical intuition. DISCUSSION: Our models produce performant, well-calibrated, and interpretable predictions for COVID-19 patients at risk for the target outcomes. They demonstrate the potential to accurately estimate outcome prognosis in resource-constrained care sites managing COVID-19 patients. CONCLUSIONS: We develop and validate prognostic models targeting MV, RRT, and readmission for hospitalized COVID-19 patients which produce accurate, interpretable predictions. Additional external validation studies are needed to further verify the generalizability of our results. Victor Alfonso Rodriguez, Shreyas Bhave, George Hripcsak, Soumitra Sengupta, Noémie Elhadad, Robert A. Green, Jason S. Adelman, Katherine Schlosser Metitiri, Pierre A. Elias, Holden Groves, Sumit Mohan, Karthik Natarajan, Adler J. Perotte |
J. Am. Medical Informatics Assoc. | 15 |
| 2021 | Clinically relevant pretraining is all you needabstractClinical notes present a wealth of information for applications in the clinical domain, but heterogeneity across clinical institutions and settings presents challenges for their processing. The clinical natural language processing field has made strides in overcoming domain heterogeneity, while pretrained deep learning models present opportunities to transfer knowledge from one task to another. Pretrained models have performed well when transferred to new tasks; however, it is not well understood if these models generalize across differences in institutions and settings within the clinical domain. We explore if institution or setting specific pretraining is necessary for pretrained models to perform well when transferred to new tasks. We find no significant performance difference between models pretrained across institutions and settings, indicating that clinically pretrained models transfer well across such boundaries. Given a clinically pretrained model, clinical natural language processing researchers may forgo the time-consuming pretraining step without a significant performance drop. Oliver J. Bear Don't Walk IV, Tony Y. Sun, Adler J. Perotte, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | A conceptual framework for external validity
Amelia J. Averitt, Patrick B. Ryan, Chunhua Weng, Adler J. Perotte |
J. Biomed. Informatics | 4 |
| 2020 | Adversarially-Learned Balancing Weights for Causal Inference
Amelia J. Averitt, Natnicha Vanitchanant, Rajesh Ranganath, Adler J. Perotte |
AMIA | 4 |
| 2020 | Deep Survival Analysis: The Impact of Feature Missingness
Shreyas Bhave, Xintian Han, Rajesh Ranganath, Adler J. Perotte |
AMIA | 4 |
| 2020 | X-CAL: Explicit Calibration for Survival AnalysisabstractSurvival analysis models the distribution of time until an event of interest, such as discharge from the hospital or admission to the ICU. When a model’s predicted number of events within any time interval is similar to the observed number, it is called well-calibrated. A survival model’s calibration can be measured using, for instance, distributional calibration (D-CALIBRATION) [Haider et al., 2020] which computes the squared difference between the observed and predicted number of events within different time intervals. Classically, calibration is addressed in post-training analysis. We develop explicit calibration (X-CAL), which turns D-CALIBRATION into a differentiable objective that can be used in survival modeling alongside maximum likelihood estimation and other objectives. X-CAL allows us to directly optimize calibration and strike a desired trade-off between predictive power and calibration. In our experiments, we fit a variety of shallow and deep models on simulated data, a survival dataset based on MNIST, on length-of-stay prediction using MIMIC-III data, and on brain cancer data from The Cancer Genome Atlas. We show that the models we study can be miscalibrated. We give experimental evidence on these datasets that X-CAL improves D-CALIBRATION without a large decrease in concordance or likelihood. Mark Goldstein, Xintian Han, Aahlad Manas Puli, Adler J. Perotte, Rajesh Ranganath |
NeurIPS | 4 |
| 2020 | Causal Estimation with Functional ConfoundersabstractCausal inference relies on two fundamental assumptions: ignorability and positivity. We study causal inference when the true confounder value can be expressed as a function of the observed data; we call this setting estimation with functional confounders (EFC). In this setting ignorability is satisfied, however positivity is violated, and causal inference is impossible in general. We consider two scenarios where causal effects are estimable. First, we discuss interventions on a part of the treatment called functional interventions and a sufficient condition for effect estimation of these interventions called functional positivity. Second, we develop conditions for nonparametric effect estimation based on the gradient fields of the functional confounder and the true outcome function. To estimate effects under these conditions, we develop Level-set Orthogonal Descent Estimation (LODE). Further, we prove error bounds on LODE’s effect estimates, evaluate our methods on simulated and real data, and empirically demonstrate the value of EFC. Aahlad Manas Puli, Adler J. Perotte, Rajesh Ranganath |
NeurIPS | 2 |
| 2020 | The Counterfactual χ-GAN: Finding comparable cohorts in observational health data
Amelia J. Averitt, Natnicha Vanitchanant, Rajesh Ranganath, Adler J. Perotte |
J. Biomed. Informatics | 4 |
| 2019 | Characterizing the Urban Opioid Epidemic Using EHR Data
Amelia J. Averitt, Benjamin H. Slovis, Abdul A. Tariq, David K. Vawdrey, Adler J. Perotte |
AMIA | 5 |
| 2019 | Learning Disease Phenotypes with Semi-Supervision
Victor Alfonso Rodriguez, Adler J. Perotte |
AMIA | 2 |
| 2019 | Challenges with quality of race and ethnicity data in observational databasesabstractOBJECTIVE: We sought to assess the quality of race and ethnicity information in observational health databases, including electronic health records (EHRs), and to propose patient self-recording as an improvement strategy. MATERIALS AND METHODS: We assessed completeness of race and ethnicity information in large observational health databases in the United States (Healthcare Cost and Utilization Project and Optum Labs), and at a single healthcare system in New York City serving a racially and ethnically diverse population. We compared race and ethnicity data collected via administrative processes with data recorded directly by respondents via paper surveys (National Health and Nutrition Examination Survey and Hospital Consumer Assessment of Healthcare Providers and Systems). Respondent-recorded data were considered the gold standard for the collection of race and ethnicity information. RESULTS: Among the 160 million patients from the Healthcare Cost and Utilization Project and Optum Labs datasets, race or ethnicity was unknown for 25%. Among the 2.4 million patients in the single New York City healthcare system's EHR, race or ethnicity was unknown for 57%. However, when patients directly recorded their race and ethnicity, 86% provided clinically meaningful information, and 66% of patients reported information that was discrepant with the EHR. DISCUSSION: Race and ethnicity data are critical to support precision medicine initiatives and to determine healthcare disparities; however, the quality of this information in observational databases is concerning. Patient self-recording through the use of patient-facing tools can substantially increase the quality of the information while engaging patients in their health. CONCLUSIONS: Patient self-recording may improve the completeness of race and ethnicity information. Fernanda Polubriaginof, Patrick B. Ryan, Hojjat Salmasian, Andrea W. Shapiro, Adler J. Perotte, Monika M. Safford, George Hripcsak, Shaun Smith, Nicholas P. Tatonetti, David K. Vawdrey |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | Noisy-Or Risk Allocation Model for Causal Inference
Amelia J. Averitt, Adler J. Perotte |
AMIA | 2 |
| 2018 | Development of a Machine Learning Model for Prediction of Successful Extubation and User-Friendly Implementation for Real-World Use
Joongheum Park, Adler J. Perotte |
AMIA | 3 |
| 2017 | Clinical Trial Eligibility Criteria Fail to Meet Burden of Generalizability
Amelia J. Averitt, Chunhua Weng, Adler J. Perotte |
AMIA | 3 |
| 2017 | Lessons Learned from the Conversion of MIMIC3 to the OHDSI Common Data Model
Joongheum Park, Liana Tascau, Adler J. Perotte |
AMIA | 3 |
| 2016 | Approaches for using temporal and other filters for next generation phenotype discovery
David J. Albers, Adler J. Perotte, George Hripcsak |
AMIA | 2 |
| 2016 | Standardization of FDA Adverse Event Reporting System to the OHDSI Common Data Model
Amelia J. Averitt, Adler J. Perotte |
AMIA | 2 |
| 2016 | A probabilistic model for learning relationships between diagnosis codes and clinical free text
Adler J. Perotte, Noémie Elhadad |
AMIA | 1 |
| 2016 | Patient-provided Data Improves Race and Ethnicity Data Quality in Electronic Health Records
Fernanda Polubriaginof, Hojjat Salmasian, Andrea W. Shapiro, Jennifer E. Prey, George Hripcsak, Adler J. Perotte, Nicholas P. Tatonetti, David K. Vawdrey |
AMIA | 6 |
| 2015 | Feasibility of Converting the Medicare Synthetic Public Use Data Into a Standardized Data Model for Clinical Research Informatics
Mark D. Danese, Erica A. Voss, Jennifer Duryea, Michelle Gleeson, Ryan Duryea, Amy Matcho, Donald O'Hara, William E. Stephens, Adler J. Perotte, Lee Evans, Christian G. Reich |
AMIA | 9 |
| 2015 | The Survival Filter: Joint Survival Analysis with a Latent Time Series
Rajesh Ranganath, Adler J. Perotte, Noémie Elhadad, David M. Blei |
UAI | 2 |
| 2015 | Parameterizing time in electronic health record studiesabstractBACKGROUND: Fields like nonlinear physics offer methods for analyzing time series, but many methods require that the time series be stationary-no change in properties over time.Objective Medicine is far from stationary, but the challenge may be able to be ameliorated by reparameterizing time because clinicians tend to measure patients more frequently when they are ill and are more likely to vary. METHODS: We compared time parameterizations, measuring variability of rate of change and magnitude of change, and looking for homogeneity of bins of temporal separation between pairs of time points. We studied four common laboratory tests drawn from 25 years of electronic health records on 4 million patients. RESULTS: We found that sequence time-that is, simply counting the number of measurements from some start-produced more stationary time series, better explained the variation in values, and had more homogeneous bins than either traditional clock time or a recently proposed intermediate parameterization. Sequence time produced more accurate predictions in a single Gaussian process model experiment. CONCLUSIONS: Of the three parameterizations, sequence time appeared to produce the most stationary series, possibly because clinicians adjust their sampling to the acuity of the patient. Parameterizing by sequence time may be applicable to association and clustering experiments on electronic health record data. A limitation of this study is that laboratory data were derived from only one institution. Sequence time appears to be an important potential parameterization. George Hripcsak, David J. Albers, Adler J. Perotte |
J. Am. Medical Informatics Assoc. | 3 |
| 2015 | Risk prediction for chronic kidney disease progression using heterogeneous electronic health record data and time series analysisabstractBACKGROUND: As adoption of electronic health records continues to increase, there is an opportunity to incorporate clinical documentation as well as laboratory values and demographics into risk prediction modeling. OBJECTIVE: The authors develop a risk prediction model for chronic kidney disease (CKD) progression from stage III to stage IV that includes longitudinal data and features drawn from clinical documentation. METHODS: The study cohort consisted of 2908 primary-care clinic patients who had at least three visits prior to January 1, 2013 and developed CKD stage III during their documented history. Development and validation cohorts were randomly selected from this cohort and the study datasets included longitudinal inpatient and outpatient data from these populations. Time series analysis (Kalman filter) and survival analysis (Cox proportional hazards) were combined to produce a range of risk models. These models were evaluated using concordance, a discriminatory statistic. RESULTS: A risk model incorporating longitudinal data on clinical documentation and laboratory test results (concordance 0.849) predicts progression from state III CKD to stage IV CKD more accurately when compared to a similar model without laboratory test results (concordance 0.733, P<.001), a model that only considers the most recent laboratory test results (concordance 0.819, P < .031) and a model based on estimated glomerular filtration rate (concordance 0.779, P < .001). CONCLUSIONS: A risk prediction model that takes longitudinal laboratory test results and clinical documentation into consideration can predict CKD progression from stage III to stage IV more accurately than three models that do not take all of these variables into consideration. Adler J. Perotte, Rajesh Ranganath, Jamie S. Hirsch, David M. Blei, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | Learning probabilistic phenotypes from heterogeneous EHR data
Rimma Perotte, Adler J. Perotte, Edouard Grave, John Angiolillo, Chris Wiggins 0001, Noémie Elhadad |
J. Biomed. Informatics | 2 |
| 2014 | Diagnosis code assignment: models and evaluation metricsabstractBACKGROUND AND OBJECTIVE: The volume of healthcare data is growing rapidly with the adoption of health information technology. We focus on automated ICD9 code assignment from discharge summary content and methods for evaluating such assignments. METHODS: We study ICD9 diagnosis codes and discharge summaries from the publicly available Multiparameter Intelligent Monitoring in Intensive Care II (MIMIC II) repository. We experiment with two coding approaches: one that treats each ICD9 code independently of each other (flat classifier), and one that leverages the hierarchical nature of ICD9 codes into its modeling (hierarchy-based classifier). We propose novel evaluation metrics, which reflect the distances among gold-standard and predicted codes and their locations in the ICD9 tree. Experimental setup, code for modeling, and evaluation scripts are made available to the research community. RESULTS: The hierarchy-based classifier outperforms the flat classifier with F-measures of 39.5% and 27.6%, respectively, when trained on 20,533 documents and tested on 2282 documents. While recall is improved at the expense of precision, our novel evaluation metrics show a more refined assessment: for instance, the hierarchy-based classifier identifies the correct sub-tree of gold-standard codes more often than the flat classifier. Error analysis reveals that gold-standard codes are not perfect, and as such the recall and precision are likely underestimated. CONCLUSIONS: Hierarchy-based classification yields better ICD9 coding than flat classification for MIMIC patients. Automated ICD9 coding is an example of a task for which data and tools can be shared and for which the research community can work together to build on shared models and advance the state of the art. Adler J. Perotte, Rimma Perotte, Karthik Natarajan, Nicole Gray Weiskopf, Frank D. Wood, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Temporal Properties of Diagnosis Code Time Series in AggregateabstractTime series are essential to health data research and data mining. We aim to study the properties of one of the more commonly available but historically unreliable types of data: administrative diagnoses in the form of the International Classification of Diseases, Ninth Revision (ICD9) codes. We use differential entropy of ICD9 code time series as a surrogate measure for disease time course and also explore Gaussian kernel smoothing to characterize the time course of diseases in a more fine-grained way. Compared to a gold standard created by a panel of clinicians, the first model classified diseases into acute and chronic groups with a receiver operating characteristic area under curve of 0.83. In the second model, several characteristic temporal profiles were observed including permanent, chronic, and acute. In addition, condition dynamics such as the refractory period for giving birth following childbirth were observed. These models demonstrate that ICD9 codes, despite well-documented concerns, contain valid and potentially valuable temporal information. Adler J. Perotte, George Hripcsak |
IEEE J. Biomed. Health Informatics | 1 |
| 2011 | Hierarchically Supervised Latent Dirichlet AllocationabstractWe introduce hierarchically supervised latent Dirichlet allocation (HSLDA), a model for hierarchically and multiply labeled bag-of-word data. Examples of such data include web pages and their placement in directories, product descriptions and associated categories from product hierarchies, and free-text clinical records and their assigned diagnosis codes. Out-of-sample label prediction is the primary goal of this work, but improved lower-dimensional representations of the bag-of-word data are also of interest. We demonstrate HSLDA on large-scale data from clinical document labeling and retail product categorization tasks. We show that leveraging the structure from hierarchical labels improves out-of-sample label prediction substantially when compared to models that do not. Adler J. Perotte, Frank D. Wood, Noémie Elhadad, Nicholas Bartlett |
NIPS | 1 |
| 2011 | Exploiting time in electronic health record correlationsabstractOBJECTIVE: To demonstrate that a large, heterogeneous clinical database can reveal fine temporal patterns in clinical associations; to illustrate several types of associations; and to ascertain the value of exploiting time. MATERIALS AND METHODS: Lagged linear correlation was calculated between seven clinical laboratory values and 30 clinical concepts extracted from resident signout notes from a 22-year, 3-million-patient database of electronic health records. Time points were interpolated, and patients were normalized to reduce inter-patient effects. RESULTS: The method revealed several types of associations with detailed temporal patterns. Definitional associations included low blood potassium preceding 'hypokalemia.' Low potassium preceding the drug spironolactone with high potassium following spironolactone exemplified intentional and physiologic associations, respectively. Counterintuitive results such as the fact that diseases appeared to follow their effects may be due to the workflow of healthcare, in which clinical findings precede the clinician's diagnosis of a disease even though the disease actually preceded the findings. Fully exploiting time by interpolating time points produced less noisy results. DISCUSSION: Electronic health records are not direct reflections of the patient state, but rather reflections of the healthcare process and the recording process. With proper techniques and understanding, and with proper incorporation of time, interpretable associations can be derived from a large clinical database. CONCLUSION: A large, heterogeneous clinical database can reveal clinical associations, time is an important feature, and care must be taken to interpret the results. George Hripcsak, David J. Albers, Adler J. Perotte |
J. Am. Medical Informatics Assoc. | 3 |
| 2009 | A Bayesian Analysis of Dynamics in Free RecallabstractWe develop a probabilistic model of human memory performance in free recall experiments. In these experiments, a subject first studies a list of words and then tries to recall them. To model these data, we draw on both previous psychological research and statistical topic models of text documents. We assume that memories are formed by assimilating the semantic meaning of studied words (represented as a distribution over topics) into a slowly changing latent context (represented in the same space). During recall, this context is reinstated and used as a cue for retrieving studied words. By conceptualizing memory retrieval as a dynamic latent variable model, we are able to use Bayesian inference to represent uncertainty and reason about the cognitive processes underlying memory. We present a particle filter algorithm for performing approximate posterior inference, and evaluate our model on the prediction of recalled words in experimental data. By specifying the model hierarchically, we are also able to capture inter-subject variability. Richard Socher, Samuel Gershman, Adler J. Perotte, Per B. Sederberg, David M. Blei, Kenneth A. Norman |
NIPS | 3 |
| 2005 | Methods for reducing interference in the Complementary Learning Systems model: Oscillating inhibition and autonomous memory rehearsal
Kenneth A. Norman, Ehren L. Newman, Adler J. Perotte |
Neural Networks | 3 |