Peter Szolovits

dblp:11/6043 · DBLP profile ↗
← Back
74ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0001-8411-6403ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 53 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 21 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Using machine learning to develop smart reflex testing protocols
abstract
OBJECTIVE: Reflex testing protocols allow clinical laboratories to perform second line diagnostic tests on existing specimens based on the results of initially ordered tests. Reflex testing can support optimal clinical laboratory test ordering and diagnosis. In current clinical practice, reflex testing typically relies on simple "if-then" rules; however, this limits the opportunities for reflex testing since most test ordering decisions involve more complexity than traditional rule-based approaches would allow. Here, using the analyte ferritin as an example, we propose an alternative machine learning-based approach to "smart" reflex testing. METHODS: Using deidentified patient data, we developed a machine learning model to predict whether a patient getting CBC testing will also have ferritin testing ordered. We evaluate applications of this model to reflex testing by assessing its performance in comparison to possible rule-based approaches. RESULTS: Our underlying machine learning models performed moderately well in predicting ferritin test ordering (AUC=0.731 in reference to actual ordering) and demonstrated promising potential to underlie key clinical applications. In contrast, none of the many traditionally framed, rule-based, hypothetical reflex protocols we evaluated offered sufficient agreement with actual ordering to be clinically feasible. Using chart review, we further demonstrated that the strategic deployment of our model could avoid important ferritin test ordering errors. CONCLUSIONS: Machine learning may provide a foundation for new types of reflex testing with enhanced benefits for clinical diagnosis.
Matthew McDermott, Anand Dighe, Peter Szolovits, Yuan Luo 0001, Jason Baron
J. Am. Medical Informatics Assoc.3
2023 An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)
abstract
Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts.
Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001
J. Am. Medical Informatics Assoc.26
2021 emrKBQA: Creating a Clinical Knowledge-Base Question Answering Dataset
Rachita Chandra, Preethi Raghavan, Jennifer J. Liang, Diwakar Mahajan, Peter Szolovits
AMIA5
2021 ATLAS: an automated association test using probabilistically linked health records with application to genetic studies
abstract
OBJECTIVE: Large amounts of health data are becoming available for biomedical research. Synthesizing information across databases may capture more comprehensive pictures of patient health and enable novel research studies. When no gold standard mappings between patient records are available, researchers may probabilistically link records from separate databases and analyze the linked data. However, previous linked data inference methods are constrained to certain linkage settings and exhibit low power. Here, we present ATLAS, an automated, flexible, and robust association testing algorithm for probabilistically linked data. MATERIALS AND METHODS: Missing variables are imputed at various thresholds using a weighted average method that propagates uncertainty from probabilistic linkage. Next, estimated effect sizes are obtained using a generalized linear model. ATLAS then conducts the threshold combination test by optimally combining P values obtained from data imputed at varying thresholds using Fisher's method and perturbation resampling. RESULTS: In simulations, ATLAS controls for type I error and exhibits high power compared to previous methods. In a real-world genetic association study, meta-analysis of ATLAS-enabled analyses on a linked cohort with analyses using an existing cohort yielded additional significant associations between rheumatoid arthritis genetic risk score and laboratory biomarkers. DISCUSSION: Weighted average imputation weathers false matches and increases contribution of true matches to mitigate linkage error-induced bias. The threshold combination test avoids arbitrarily choosing a threshold to rule a match, thus automating linked data-enabled analyses and preserving power. CONCLUSION: ATLAS promises to enable novel and powerful research studies using linked data to capitalize on all available data sources.
Harrison G. Zhang, Boris P. Hejblum, Griffin M. Weber, Nathan P. Palmer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Katherine P. Liao, Isaac S. Kohane, Tianxi Cai
J. Am. Medical Informatics Assoc.6
2020 Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
abstract
Machine learning algorithms are often vulnerable to adversarial examples that have imperceptible alterations from the original counterparts but can fool the state-of-the-art models. It is helpful to evaluate or even improve the robustness of these models by exposing the maliciously crafted adversarial examples. In this paper, we present TextFooler, a simple but strong baseline to generate adversarial text. By applying it to two fundamental natural language tasks, text classification and textual entailment, we successfully attacked three target models, including the powerful pre-trained BERT, and the widely used convolutional and recurrent neural networks. We demonstrate three advantages of this framework: (1) effective—it outperforms previous attacks by success rate and perturbation rate, (2) utility-preserving—it preserves semantic content, grammaticality, and correct types classified by humans, and (3) efficient—it generates adversarial text with computational complexity linear to the text length.1
Di Jin 0005, Zhijing Jin 0001, Joey Tianyi Zhou, Peter Szolovits
AAAI4
2020 Hooks in the Headline: Learning to Generate Headlines with Controlled Styles
abstract
Current summarization systems only produce plain, factual headlines, but do not meet the practical needs of creating memorable titles to increase exposure.We propose a new task, Stylistic Headline Generation (SHG), to enrich the headlines with three style options (humor, romance and clickbait), in order to attract more readers.With no style-specific article-headline pair (only a standard headline summarization dataset and mono-style corpora), our method TitleStylist generates style-specific headlines by combining the summarization and reconstruction tasks into a multitasking framework.We also introduced a novel parameter sharing scheme to further disentangle the style from the text.Through both automatic and human evaluation, we demonstrate that TitleStylist can generate relevant, fluent headlines with three target styles: humor, romance, and clickbait.The attraction score of our model generated headlines surpasses that of the state-ofthe-art summarization model by 9.68%, and even outperforms human-written references. 1
Di Jin 0005, Zhijing Jin 0001, Joey Tianyi Zhou, Lisa Orii, Peter Szolovits
ACL5
2020 Joint Modeling of Chest Radiographs and Radiology Reports for Pulmonary Edema Assessment
Geeticka Chauhan, Ruizhi Liao 0001, William M. Wells III, Jacob Andreas, Seth J. Berkowitz, Steven Horng, Peter Szolovits, Polina Golland
MICCAI (2)8
2020 Expert-Supervised Reinforcement Learning for Offline Policy Learning and Evaluation
abstract
Offline Reinforcement Learning (RL) is a promising approach for learning optimal policies in environments where direct exploration is expensive or unfeasible. However, the adoption of such policies in practice is often challenging, as they are hard to interpret within the application context, and lack measures of uncertainty for the learned policy value and its decisions. To overcome these issues, we propose an Expert-Supervised RL (ESRL) framework which uses uncertainty quantification for offline policy learning. In particular, we have three contributions: 1) the method can learn safe and optimal policies through hypothesis testing, 2) ESRL allows for different levels of risk averse implementations tailored to the application context, and finally, 3) we propose a way to interpret ESRL’s policy at every state through posterior distributions, and use this framework to compute off-policy value function posteriors. We provide theoretical guarantees for our estimators and regret bounds consistent with Posterior Sampling for RL (PSRL). Sample efficiency of ESRL is independent of the chosen risk aversion threshold and quality of the behavior policy.
Aaron Sonabend W., Leo A. Celi, Tianxi Cai, Peter Szolovits
NeurIPS5
2020 Advancing PICO element detection in biomedical text via deep neural networks
abstract
MOTIVATION: In evidence-based medicine, defining a clinical question in terms of the specific patient problem aids the physicians to efficiently identify appropriate resources and search for the best available evidence for medical treatment. In order to formulate a well-defined, focused clinical question, a framework called PICO is widely used, which identifies the sentences in a given medical text that belong to the four components typically reported in clinical trials: Participants/Problem (P), Intervention (I), Comparison (C) and Outcome (O). In this work, we propose a novel deep learning model for recognizing PICO elements in biomedical abstracts. Based on the previous state-of-the-art bidirectional long-short-term memory (bi-LSTM) plus conditional random field architecture, we add another layer of bi-LSTM upon the sentence representation vectors so that the contextual information from surrounding sentences can be gathered to help infer the interpretation of the current one. In addition, we propose two methods to further generalize and improve the model: adversarial training and unsupervised pre-training over large corpora. RESULTS: We tested our proposed approach over two benchmark datasets. One is the PubMed-PICO dataset, where our best results outperform the previous best by 5.5%, 7.9% and 5.8% for P, I and O elements in terms of F1 score, respectively. And for the other dataset named NICTA-PIBOSO, the improvements for P/I/O elements are 3.9%, 15.6% and 1.3% in F1 score, respectively. Overall, our proposed deep learning model can obtain unprecedented PICO element detection accuracy while avoiding the need for any manual feature selection. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/jind11/Deep-PICO-Detection.
Di Jin 0005, Peter Szolovits
Bioinform.2
2020 Deep Learning Benchmarks on L1000 Gene Expression Data
abstract
Gene expression data can offer deep, physiological insights beyond the static coding of the genome alone. We believe that realizing this potential requires specialized, high-capacity machine learning methods capable of using underlying biological structure, but the development of such models is hampered by the lack of published benchmark tasks and well characterized baselines. In this work, we establish such benchmarks and baselines by profiling many classifiers against biologically motivated tasks on two curated views of a large, public gene expression dataset (the LINCS corpus) and one privately produced dataset. We provide these two curated views of the public LINCS dataset and our benchmark tasks to enable direct comparisons to future methodological work and help spur deep learning method development on this modality. In addition to profiling a battery of traditional classifiers, including linear models, random forests, decision trees, K nearest neighbor (KNN) classifiers, and feed-forward artificial neural networks (FF-ANNs), we also test a method novel to this data modality: graph convolugtional neural networks (GCNNs), which allow us to incorporate prior biological domain knowledge. We find that GCNNs can be highly performant, with large datasets, whereas FF-ANNs consistently perform well. Non-neural classifiers are dominated by linear models and KNN classifiers.
Matthew B. A. McDermott, Wen-Ning Zhao, Steven Sheridan, Peter Szolovits, Isaac S. Kohane, Stephen J. Haggarty, Roy H. Perlis
IEEE ACM Trans. Comput. Biol. Bioinform.5
2019 Unsupervised Clinical Language Translation
abstract
As patients' access to their doctors' clinical notes becomes common, translating professional, clinical jargon to layperson-understandable language is essential to improve patient-clinician communication. Such translation yields better clinical outcomes by enhancing patients' understanding of their own health conditions, and thus improving patients' involvement in their own care. Existing research has used dictionary-based word replacement or definition insertion to approach the need. However, these methods are limited by expert curation, which is hard to scale and has trouble generalizing to unseen datasets that do not share an overlapping vocabulary. In contrast, we approach the clinical word and sentence translation problem in a completely unsupervised manner. We show that a framework using representation learning, bilingual dictionary induction and statistical machine translation yields the best precision at 10 of 0.827 on professional-to-consumer word translation, and mean opinion scores of 4.10 and 4.28 out of 5 for clinical correctness and layperson readability, respectively, on sentence translation. Our fully-unsupervised strategy overcomes the curation problem, and the clinically meaningful evaluation reduces biases from inappropriate evaluators, which are critical in clinical machine learning.
Wei-Hung Weng, Yu-An Chung, Peter Szolovits
KDD3
2019 High-throughput multimodal automated phenotyping (MAP) with application to PheWAS
abstract
OBJECTIVE: Electronic health records linked with biorepositories are a powerful platform for translational studies. A major bottleneck exists in the ability to phenotype patients accurately and efficiently. The objective of this study was to develop an automated high-throughput phenotyping method integrating International Classification of Diseases (ICD) codes and narrative data extracted using natural language processing (NLP). MATERIALS AND METHODS: We developed a mapping method for automatically identifying relevant ICD and NLP concepts for a specific phenotype leveraging the Unified Medical Language System. Along with health care utilization, aggregated ICD and NLP counts were jointly analyzed by fitting an ensemble of latent mixture models. The multimodal automated phenotyping (MAP) algorithm yields a predicted probability of phenotype for each patient and a threshold for classifying participants with phenotype yes/no. The algorithm was validated using labeled data for 16 phenotypes from a biorepository and further tested in an independent cohort phenome-wide association studies (PheWAS) for 2 single nucleotide polymorphisms with known associations. RESULTS: The MAP algorithm achieved higher or similar AUC and F-scores compared to the ICD code across all 16 phenotypes. The features assembled via the automated approach had comparable accuracy to those assembled via manual curation (AUCMAP 0.943, AUCmanual 0.941). The PheWAS results suggest that the MAP approach detected previously validated associations with higher power when compared to the standard PheWAS method based on ICD codes. CONCLUSION: The MAP approach increased the accuracy of phenotype definition while maintaining scalability, thereby facilitating use in studies requiring large-scale phenotyping, such as PheWAS.
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Yuk-Lam Ho, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Christopher J. O'Donnell, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002
J. Am. Medical Informatics Assoc.17
2018 Semi-Supervised Biomedical Translation With Cycle Wasserstein Regression GANs
abstract
The biomedical field offers many learning tasks that share unique challenges: large amounts of unpaired data, and a high cost to generate labels. In this work, we develop a method to address these issues with semi-supervised learning in regression tasks (e.g., translation from source to target). Our model uses adversarial signals to learn from unpaired datapoints, and imposes a cycle-loss reconstruction error penalty to regularize mappings in either direction against one another. We first evaluate our method on synthetic experiments, demonstrating two primary advantages of the system: 1) distribution matching via the adversarial loss and 2) regularization towards invertible mappings via the cycle loss. We then show a regularization effect and improved performance when paired data is supplemented by additional unpaired data on two real biomedical regression tasks: estimating the physiological effect of medical treatments, and extrapolating gene expression (transcriptomics) signals. Our proposed technique is a promising initial step towards more robust use of adversarial signals in semi-supervised regression, and could be useful for other tasks (e.g., causal inference or modality translation) in the biomedical field.
Matthew B. A. McDermott, Tom Yan, Tristan Naumann, Nathan Hunt, Harini Suresh, Peter Szolovits, Marzyeh Ghassemi
AAAI6
2018 High-Throughput Multimodal Automated Phenotyping (MAP) Incorporating Natural Language Processing with Application to PheWAS
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Lauren Costa, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002, Tianxi Cai
AMIA15
2018 Implementing a Portable Clinical NLP System with a Common Data Model - a Lisp Perspective
Yuan Luo 0001, Peter Szolovits
BIBM2
2018 Hierarchical Neural Networks for Sequential Sentence Classification in Medical Scientific Abstracts
abstract
Prevalent models based on artificial neural network (ANN) for sentence classification often classify sentences in isolation without considering the context in which sentences appear.This hampers the traditional sentence classification approaches to the problem of sequential sentence classification, where structured prediction is needed for better overall classification performance.In this work, we present a hierarchical sequential labeling network to make use of the contextual information within surrounding sentences to help classify the current sentence.Our model outperforms the state-of-the-art results by 2%-3% on two benchmarking datasets for sequential sentence classification in medical scientific abstracts.
Di Jin 0005, Peter Szolovits
EMNLP2
2018 Transfer Learning for Named-Entity Recognition with Neural Networks
Ji Young Lee 0001, Franck Dernoncourt, Peter Szolovits
LREC3
2018 Segment convolutional neural networks (Seg-CNNs) for classifying relations in clinical notes
abstract
We propose Segment Convolutional Neural Networks (Seg-CNNs) for classifying relations from clinical notes. Seg-CNNs use only word-embedding features without manual feature engineering. Unlike typical CNN models, relations between 2 concepts are identified by simultaneously learning separate representations for text segments in a sentence: preceding, concept1, middle, concept2, and succeeding. We evaluate Seg-CNN on the i2b2/VA relation classification challenge dataset. We show that Seg-CNN achieves a state-of-the-art micro-average F-measure of 0.742 for overall evaluation, 0.686 for classifying medical problem-treatment relations, 0.820 for medical problem-test relations, and 0.702 for medical problem-medical problem relations. We demonstrate the benefits of learning segment-level representations. We show that medical domain word embeddings help improve relation classification. Seg-CNNs can be trained quickly for the i2b2/VA dataset on a graphics processing unit (GPU) platform. These results support the use of CNNs computed over segments of text for classifying medical relations, as they show state-of-the-art performance while requiring no manual feature engineering.
Yuan Luo 0001, Özlem Uzuner, Peter Szolovits, Justin Starren
J. Am. Medical Informatics Assoc.4
2018 3D-MICE: integration of cross-sectional and longitudinal imputation for multi-analyte longitudinal clinical data
abstract
Objective: A key challenge in clinical data mining is that most clinical datasets contain missing data. Since many commonly used machine learning algorithms require complete datasets (no missing data), clinical analytic approaches often entail an imputation procedure to "fill in" missing data. However, although most clinical datasets contain a temporal component, most commonly used imputation methods do not adequately accommodate longitudinal time-based data. We sought to develop a new imputation algorithm, 3-dimensional multiple imputation with chained equations (3D-MICE), that can perform accurate imputation of missing clinical time series data. Methods: We extracted clinical laboratory test results for 13 commonly measured analytes (clinical laboratory tests). We imputed missing test results for the 13 analytes using 3 imputation methods: multiple imputation with chained equations (MICE), Gaussian process (GP), and 3D-MICE. 3D-MICE utilizes both MICE and GP imputation to integrate cross-sectional and longitudinal information. To evaluate imputation method performance, we randomly masked selected test results and imputed these masked results alongside results missing from our original data. We compared predicted results to measured results for masked data points. Results: 3D-MICE performed significantly better than MICE and GP-based imputation in a composite of all 13 analytes, predicting missing results with a normalized root-mean-square error of 0.342, compared to 0.373 for MICE alone and 0.358 for GP alone. Conclusions: 3D-MICE offers a novel and practical approach to imputing clinical laboratory time series data. 3D-MICE may provide an additional tool for use as a foundation in clinical predictive analytics and intelligent clinical decision support.
Yuan Luo 0001, Peter Szolovits, Anand Dighe, Jason Baron
J. Am. Medical Informatics Assoc.2
2018 Enabling phenotypic big data with PheNorm
abstract
Objective: Electronic health record (EHR)-based phenotyping infers whether a patient has a disease based on the information in his or her EHR. A human-annotated training set with gold-standard disease status labels is usually required to build an algorithm for phenotyping based on a set of predictive features. The time intensiveness of annotation and feature curation severely limits the ability to achieve high-throughput phenotyping. While previous studies have successfully automated feature curation, annotation remains a major bottleneck. In this paper, we present PheNorm, a phenotyping algorithm that does not require expert-labeled samples for training. Methods: The most predictive features, such as the number of International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) codes or mentions of the target phenotype, are normalized to resemble a normal mixture distribution with high area under the receiver operating curve (AUC) for prediction. The transformed features are then denoised and combined into a score for accurate disease classification. Results: We validated the accuracy of PheNorm with 4 phenotypes: coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis. The AUCs of the PheNorm score reached 0.90, 0.94, 0.95, and 0.94 for the 4 phenotypes, respectively, which were comparable to the accuracy of supervised algorithms trained with sample sizes of 100-300, with no statistically significant difference. Conclusion: The accuracy of the PheNorm algorithms is on par with algorithms trained with annotated samples. PheNorm fully automates the generation of accurate phenotyping algorithms and demonstrates the capacity for EHR-driven annotations to scale to the next level - phenotypic big data.
Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai
J. Am. Medical Informatics Assoc.8
2017 High-throughput Phenotyping via Denoised Normal Mixture Transformation
Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai
AMIA9
2017 Predicting Clinical Outcomes Across Changing Electronic Health Record Systems
abstract
Existing machine learning methods typically assume consistency in how semantically equivalent information is encoded. However, the way information is recorded in databases differs across institutions and over time, often rendering potentially useful data obsolescent. To address this problem, we map database-specific representations of information to a shared set of semantic concepts, thus allowing models to be built from or transition across different databases. We demonstrate our method on machine learning models developed in a healthcare setting. In particular, we evaluate our method using two different intensive care unit (ICU) databases and on two clinically relevant tasks, in-hospital mortality and prolonged length of stay. For both outcomes, a feature representation mapping EHR-specific events to a shared set of clinical concepts yields better results than using EHR-specific events alone.
Jen J. Gong, Tristan Naumann, Peter Szolovits, John V. Guttag
KDD3
2017 Bridging semantics and syntax with graph algorithms - state-of-the-art of extracting biomedical relations
abstract
Research on extracting biomedical relations has received growing attention recently, with numerous biological and clinical applications including those in pharmacogenomics, clinical trial screening and adverse drug reaction detection. The ability to accurately capture both semantic and syntactic structures in text expressing these relations becomes increasingly critical to enable deep understanding of scientific papers and clinical narratives. Shared task challenges have been organized by both bioinformatics and clinical informatics communities to assess and advance the state-of-the-art research. Significant progress has been made in algorithm development and resource construction. In particular, graph-based approaches bridge semantics and syntax, often achieving the best performance in shared tasks. However, a number of problems at the frontiers of biomedical relation extraction continue to pose interesting challenges and present opportunities for great improvement and fruitful research. In this article, we place biomedical relation extraction against the backdrop of its versatile applications, present a gentle introduction to its general pipeline and shared resources, review the current state-of-the-art in methodology advancement, discuss limitations and point out several promising future directions.
Yuan Luo 0001, Özlem Uzuner, Peter Szolovits
Briefings Bioinform.3
2017 Bridging semantics and syntax with graph algorithms - state-of-the-art of extracting biomedical relations
abstract
Briefings in Bioinformatics (2017) 18(1), 2017, 160–178, doi: 10.1093/bib/bbw001 In the above article, the sentence ‘Wang et al. [84] used Latent Dirichl et al. location to create a semantic representation of biomedical named entities and used Kullback-Leibler (KL) divergence to calculate the association distance between pairs of entities in the Chem2Bio2RDF [149] semantic network’ has been corrected to ‘Wang et al. [84] used Latent Dirichlet Allocation to create a semantic representation of biomedical named entities and used Kullback-Leibler (KL) divergence to calculate the association distance between pairs of entities in the Chem2Bio2RDF [149] semantic network’. The text has been corrected online. The publisher apologizes for this error.
Yuan Luo 0001, Özlem Uzuner, Peter Szolovits
Briefings Bioinform.3
2017 Tensor factorization toward precision medicine
abstract
Precision medicine initiatives come amid the rapid growth in quantity and variety of biomedical data, which exceeds the capacity of matrix-oriented data representations and many current analysis algorithms. Tensor factorizations extend the matrix view to multiple modalities and support dimensionality reduction methods that identify latent groups of data for meaningful summarization of both features and instances. In this opinion article, we analyze the modest literature on applying tensor factorization to various biomedical fields including genotyping and phenotyping. Based on the cited work including work of our own, we suggest that tensor applications could serve as an effective tool to enable frequent updating of medical knowledge based on the continually growing scientific and clinical evidence. We encourage extensive experimental studies to tackle challenges including design choice of factorizations, integrating temporality and algorithm scalability.
Yuan Luo 0001, Fei Wang 0001, Peter Szolovits
Briefings Bioinform.3
2017 De-identification of patient notes with recurrent neural networks
abstract
OBJECTIVE: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the confidentiality of patients. In the United States, the Health Insurance Portability and Accountability Act (HIPAA) defines 18 types of protected health information that needs to be removed to de-identify patient notes. Manual de-identification is impractical given the size of electronic health record databases, the limited number of researchers with access to non-de-identified notes, and the frequent mistakes of human annotators. A reliable automated de-identification system would consequently be of high value. MATERIALS AND METHODS: We introduce the first de-identification system based on artificial neural networks (ANNs), which requires no handcrafted features or rules, unlike existing systems. We compare the performance of the system with state-of-the-art systems on two datasets: the i2b2 2014 de-identification challenge dataset, which is the largest publicly available de-identification dataset, and the MIMIC de-identification dataset, which we assembled and is twice as large as the i2b2 2014 dataset. RESULTS: Our ANN model outperforms the state-of-the-art systems. It yields an F1-score of 97.85 on the i2b2 2014 dataset, with a recall of 97.38 and a precision of 98.32, and an F1-score of 99.23 on the MIMIC de-identification dataset, with a recall of 99.25 and a precision of 99.21. CONCLUSION: Our findings support the use of ANNs for de-identification of patient notes, as they show better performance than previously published systems while requiring no manual feature engineering.
Franck Dernoncourt, Ji Young Lee 0001, Özlem Uzuner, Peter Szolovits
J. Am. Medical Informatics Assoc.4
2017 Understanding vasopressor intervention and weaning: risk prediction in a public heterogeneous clinical time series database
abstract
BACKGROUND: The widespread adoption of electronic health records allows us to ask evidence-based questions about the need for and benefits of specific clinical interventions in critical-care settings across large populations. OBJECTIVE: We investigated the prediction of vasopressor administration and weaning in the intensive care unit. Vasopressors are commonly used to control hypotension, and changes in timing and dosage can have a large impact on patient outcomes. MATERIALS AND METHODS: We considered a cohort of 15 695 intensive care unit patients without orders for reduced care who were alive 30 days post-discharge. A switching-state autoregressive model (SSAM) was trained to predict the multidimensional physiological time series of patients before, during, and after vasopressor administration. The latent states from the SSAM were used as predictors of vasopressor administration and weaning. RESULTS: The unsupervised SSAM features were able to predict patient vasopressor administration and successful patient weaning. Features derived from the SSAM achieved areas under the receiver operating curve of 0.92, 0.88, and 0.71 for predicting ungapped vasopressor administration, gapped vasopressor administration, and vasopressor weaning, respectively. We also demonstrated many cases where our model predicted weaning well in advance of a successful wean. CONCLUSION: Models that used SSAM features increased performance on both predictive tasks. These improvements may reflect an underlying, and ultimately predictive, latent state detectable from the physiological time series.
Mike Wu, Marzyeh Ghassemi, Mengling Feng, Leo A. Celi, Peter Szolovits, Finale Doshi-Velez
J. Am. Medical Informatics Assoc.5
2017 Surrogate-assisted feature extraction for high-throughput phenotyping
abstract
OBJECTIVE: Phenotyping algorithms are capable of accurately identifying patients with specific phenotypes from within electronic medical records systems. However, developing phenotyping algorithms in a scalable way remains a challenge due to the extensive human resources required. This paper introduces a high-throughput unsupervised feature selection method, which improves the robustness and scalability of electronic medical record phenotyping without compromising its accuracy. METHODS: The proposed Surrogate-Assisted Feature Extraction (SAFE) method selects candidate features from a pool of comprehensive medical concepts found in publicly available knowledge sources. The target phenotype's International Classification of Diseases, Ninth Revision and natural language processing counts, acting as noisy surrogates to the gold-standard labels, are used to create silver-standard labels. Candidate features highly predictive of the silver-standard labels are selected as the final features. RESULTS: Algorithms were trained to identify patients with coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis using various numbers of labels to compare the performance of features selected by SAFE, a previously published automated feature extraction for phenotyping procedure, and domain experts. The out-of-sample area under the receiver operating characteristic curve and F -score from SAFE algorithms were remarkably higher than those from the other two, especially at small label sizes. CONCLUSION: SAFE advances high-throughput phenotyping methods by automatically selecting a succinct set of informative features for algorithm training, which in turn reduces overfitting and the needed number of gold-standard labels. SAFE also potentially identifies important features missed by automated feature extraction for phenotyping or experts.
Sheng Yu 0002, Abhishek Chakrabortty, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai
J. Am. Medical Informatics Assoc.8
2017 Predicting Social Anxiety Treatment Outcome Based on Therapeutic Email Conversations
abstract
Predicting therapeutic outcome in the mental health domain is of utmost importance to enable therapists to provide the most effective treatment to a patient. Using information from the writings of a patient can potentially be a valuable source of information, especially now that more and more treatments involve computer-based exercises or electronic conversations between patient and therapist. In this paper, we study predictive modeling using writings of patients under treatment for a social anxiety disorder. We extract a wealth of information from the text written by patients including their usage of words, the topics they talk about, the sentiment of the messages, and the style of writing. In addition, we study trends over time with respect to those measures. We then apply machine learning algorithms to generate the predictive models. Based on a dataset of 69 patients, we are able to show that we can predict therapy outcome with an area under the curve of 0.83 halfway through the therapy and with a precision of 0.78 when using the full data (i.e., the entire treatment period). Due to the limited number of participants, it is hard to generalize the results, but they do show great potential in this type of information.
Mark Hoogendoorn, Thomas Berger 0003, Ava Schulz, Timo Stolz, Peter Szolovits
IEEE J. Biomed. Health Informatics5
2016 Predicting ICU Mortality Risk by Grouping Temporal Trends from a Multivariate Panel of Physiologic Measurements
abstract
ICU mortality risk prediction may help clinicians take effective interventions to improve patient outcome. Existing machine learning approaches often face challenges in integrating a comprehensive panel of physiologic variables and presenting to clinicians interpretable models. We aim to improve both accuracy and interpretability of prediction models by introducing Subgraph Augmented Non-negative Matrix Factorization (SANMF) on ICU physiologic time series. SANMF converts time series into a graph representation and applies frequent subgraph mining to automatically extract temporal trends. We then apply non-negative matrix factorization to group trends in a way that approximates patient pathophysiologic states. Trend groups are then used as features in training a logistic regression model for mortality risk prediction, and are also ranked according to their contribution to mortality risk. We evaluated SANMF against four empirical models on the task of predicting mortality or survival 30 days after discharge from ICU using the observed physiologic measurements between 12 and 24 hours after admission. SANMF outperforms all comparison models, and in particular, demonstrates an improvement in AUC (0.848 vs. 0.827, p<0.002) compared to a state-of-the-art machine learning method that uses manual feature engineering. Feature analysis was performed to illuminate insights and benefits of subgraph groups in mortality risk prediction.
Yuan Luo 0001, Rohit Joshi, Leo A. Celi, Peter Szolovits
AAAI5
2016 Utilizing uncoded consultation notes from electronic medical records for predictive modeling of colorectal cancer
Mark Hoogendoorn, Peter Szolovits, Leon M. G. Moons, Mattijs E. Numans
Artif. Intell. Medicine2
2015 A Multivariate Timeseries Modeling Approach to Severity of Illness Assessment and Forecasting in ICU with Sparse, Heterogeneous Clinical Data
abstract
The ability to determine patient acuity (or severity of illness) has immediate practical use for clinicians. We evaluate the use of multivariate timeseries modeling with the multi-task Gaussian process (GP) models using noisy, incomplete, sparse, heterogeneous and unevenly-sampled clinical data, including both physiological signals and clinical notes. The learned multi-task GP (MTGP) hyperparameters are then used to assess and forecast patient acuity. Experiments were conducted with two real clinical data sets acquired from ICU patients: firstly, estimating cerebrovascular pressure reactivity, an important indicator of secondary damage for traumatic brain injury patients, by learning the interactions between intracranial pressure and mean arterial blood pressure signals, and secondly, mortality prediction using clinical progress notes. In both cases, MTGPs provided improved results: an MTGP model provided better results than single-task GP models for signal interpolation and forecasting (0.91 vs 0.69 RMSE), and the use of MTGP hyperparameters obtained improved results when used as additional classification features (0.812 vs 0.788 AUC).
Marzyeh Ghassemi, Marco A. F. Pimentel, Tristan Naumann, Thomas Brennan, David A. Clifton, Peter Szolovits, Mengling Feng
AAAI6
2015 Demonstrating the Advantages of Applying Data Mining Techniques on Time-Dependent Electronic Medical Records
Uri Kartoun, Vishesh Kumar, Su-Chun Cheng, Sheng Yu 0002, Katherine P. Liao, Elizabeth W. Karlson, Ashwin N. Ananthakrishnan, Zongqi Xia, Vivian S. Gainer, Andrew Cagan, Guergana K. Savova, Pei J. Chen, Shawn N. Murphy, Susanne E. Churchill, Isaac S. Kohane, Peter Szolovits, Tianxi Cai, Stanley Y. Shaw
AMIA16
2015 Subgraph augmented non-negative tensor factorization (SANTF) for modeling clinical narrative text
abstract
OBJECTIVE: Extracting medical knowledge from electronic medical records requires automated approaches to combat scalability limitations and selection biases. However, existing machine learning approaches are often regarded by clinicians as black boxes. Moreover, training data for these automated approaches at often sparsely annotated at best. The authors target unsupervised learning for modeling clinical narrative text, aiming at improving both accuracy and interpretability. METHODS: The authors introduce a novel framework named subgraph augmented non-negative tensor factorization (SANTF). In addition to relying on atomic features (e.g., words in clinical narrative text), SANTF automatically mines higher-order features (e.g., relations of lymphoid cells expressing antigens) from clinical narrative text by converting sentences into a graph representation and identifying important subgraphs. The authors compose a tensor using patients, higher-order features, and atomic features as its respective modes. We then apply non-negative tensor factorization to cluster patients, and simultaneously identify latent groups of higher-order features that link to patient clusters, as in clinical guidelines where a panel of immunophenotypic features and laboratory results are used to specify diagnostic criteria. RESULTS AND CONCLUSION: SANTF demonstrated over 10% improvement in averaged F-measure on patient clustering compared to widely used non-negative matrix factorization (NMF) and k-means clustering methods. Multiple baselines were established by modeling patient data using patient-by-features matrices with different feature configurations and then performing NMF or k-means to cluster patients. Feature analysis identified latent groups of higher-order features that lead to medical insights. We also found that the latent groups of atomic features help to better correlate the latent groups of higher-order features.
Yuan Luo 0001, Ephraim P. Hochberg, Rohit Joshi, Özlem Uzuner, Peter Szolovits
J. Am. Medical Informatics Assoc.6
2015 Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sources
abstract
OBJECTIVE: Analysis of narrative (text) data from electronic health records (EHRs) can improve population-scale phenotyping for clinical and genetic research. Currently, selection of text features for phenotyping algorithms is slow and laborious, requiring extensive and iterative involvement by domain experts. This paper introduces a method to develop phenotyping algorithms in an unbiased manner by automatically extracting and selecting informative features, which can be comparable to expert-curated ones in classification accuracy. MATERIALS AND METHODS: Comprehensive medical concepts were collected from publicly available knowledge sources in an automated, unbiased fashion. Natural language processing (NLP) revealed the occurrence patterns of these concepts in EHR narrative notes, which enabled selection of informative features for phenotype classification. When combined with additional codified features, a penalized logistic regression model was trained to classify the target phenotype. RESULTS: The authors applied our method to develop algorithms to identify patients with rheumatoid arthritis and coronary artery disease cases among those with rheumatoid arthritis from a large multi-institutional EHR. The area under the receiver operating characteristic curves (AUC) for classifying RA and CAD using models trained with automated features were 0.951 and 0.929, respectively, compared to the AUCs of 0.938 and 0.929 by models trained with expert-curated features. DISCUSSION: Models trained with NLP text features selected through an unbiased, automated procedure achieved comparable or slightly higher accuracy than those trained with expert-curated features. The majority of the selected model features were interpretable. CONCLUSION: The proposed automated feature extraction method, generating highly accurate phenotyping algorithms with improved efficiency, is a significant step toward high-throughput phenotyping.
Sheng Yu 0002, Katherine P. Liao, Stanley Y. Shaw, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai
J. Am. Medical Informatics Assoc.6
2014 Quantifying Information Redundancy in Common Laboratory Tests
Yuan Luo 0001, Jason Baron, Peter Szolovits, Anand Dighe
AMIA3
2014 Unfolding physiological state: mortality modelling in intensive care units
abstract
Accurate knowledge of a patient's disease state and trajectory is critical in a clinical setting. Modern electronic healthcare records contain an increasingly large amount of data, and the ability to automatically identify the factors that influence patient outcomes stand to greatly improve the efficiency and quality of care. We examined the use of latent variable models (viz. Latent Dirichlet Allocation) to decompose free-text hospital notes into meaningful features, and the predictive power of these features for patient mortality. We considered three prediction regimes: (1) baseline prediction, (2) dynamic (time-varying) outcome prediction, and (3) retrospective outcome prediction. In each, our prediction task differs from the familiar time-varying situation whereby data accumulates; since fewer patients have long ICU stays, as we move forward in time fewer patients are available and the prediction task becomes increasingly difficult. We found that latent topic-derived features were effective in determining patient mortality under three timelines: inhospital, 30 day post-discharge, and 1 year post-discharge mortality. Our results demonstrated that the latent topic features important in predicting hospital mortality are very different from those that are important in post-discharge mortality. In general, latent topic features were more predictive than structured features, and a combination of the two performed best. The time-varying models that combined latent topic features and baseline features had AUCs that reached 0.85, 0.80, and 0.77 for in-hospital, 30 day post-discharge and 1 year post-discharge mortality respectively. Our results agreed with other work suggesting that the first 24 hours of patient information are often the most predictive of hospital mortality. Retrospective models that used a combination of latent topic features and structured features achieved AUCs of 0.96, 0.82, and 0.81 for in-hospital, 30 day, and 1-year mortality prediction. Our work focuses on the dynamic (time-varying) setting because models from this regime could facilitate an on-going severity stratification system that helps direct care-staff resources and inform treatment strategies.
Marzyeh Ghassemi, Tristan Naumann, Finale Doshi-Velez, Nicole Brimmer, Rohit Joshi, Anna Rumshisky, Peter Szolovits
KDD7
2014 Research and applications: Word sense disambiguation in the clinical domain: a comparison of knowledge-rich and knowledge-poor unsupervised methods
abstract
OBJECTIVE: To evaluate state-of-the-art unsupervised methods on the word sense disambiguation (WSD) task in the clinical domain. In particular, to compare graph-based approaches relying on a clinical knowledge base with bottom-up topic-modeling-based approaches. We investigate several enhancements to the topic-modeling techniques that use domain-specific knowledge sources. MATERIALS AND METHODS: The graph-based methods use variations of PageRank and distance-based similarity metrics, operating over the Unified Medical Language System (UMLS). Topic-modeling methods use unlabeled data from the Multiparameter Intelligent Monitoring in Intensive Care (MIMIC II) database to derive models for each ambiguous word. We investigate the impact of using different linguistic features for topic models, including UMLS-based and syntactic features. We use a sense-tagged clinical dataset from the Mayo Clinic for evaluation. RESULTS: The topic-modeling methods achieve 66.9% accuracy on a subset of the Mayo Clinic's data, while the graph-based methods only reach the 40-50% range, with a most-frequent-sense baseline of 56.5%. Features derived from the UMLS semantic type and concept hierarchies do not produce a gain over bag-of-words features in the topic models, but identifying phrases from UMLS and using syntax does help. DISCUSSION: Although topic models outperform graph-based methods, semantic features derived from the UMLS prove too noisy to improve performance beyond bag-of-words. CONCLUSIONS: Topic modeling for WSD provides superior results in the clinical domain; however, integration of knowledge remains to be effectively exploited.
Rachel Chasin, Anna Rumshisky, Özlem Uzuner, Peter Szolovits
J. Am. Medical Informatics Assoc.4
2014 Research and applications: Automatic lymphoma classification with sentence subgraph mining from pathology reports
abstract
OBJECTIVE: Pathology reports are rich in narrative statements that encode a complex web of relations among medical concepts. These relations are routinely used by doctors to reason on diagnoses, but often require hand-crafted rules or supervised learning to extract into prespecified forms for computational disease modeling. We aim to automatically capture relations from narrative text without supervision. METHODS: We design a novel framework that translates sentences into graph representations, automatically mines sentence subgraphs, reduces redundancy in mined subgraphs, and automatically generates subgraph features for subsequent classification tasks. To ensure meaningful interpretations over the sentence graphs, we use the Unified Medical Language System Metathesaurus to map token subsequences to concepts, and in turn sentence graph nodes. We test our system with multiple lymphoma classification tasks that together mimic the differential diagnosis by a pathologist. To this end, we prevent our classifiers from looking at explicit mentions or synonyms of lymphomas in the text. RESULTS AND CONCLUSIONS: We compare our system with three baseline classifiers using standard n-grams, full MetaMap concepts, and filtered MetaMap concepts. Our system achieves high F-measures on multiple binary classifications of lymphoma (Burkitt lymphoma, 0.8; diffuse large B-cell lymphoma, 0.909; follicular lymphoma, 0.84; Hodgkin lymphoma, 0.912). Significance tests show that our system outperforms all three baselines. Moreover, feature analysis identifies subgraph features that contribute to improved performance; these features agree with the state-of-the-art knowledge about lymphoma classification. We also highlight how these unsupervised relation features may provide meaningful insights into lymphoma classification.
Yuan Luo 0001, Aliyah R. Sohani, Ephraim P. Hochberg, Peter Szolovits
J. Am. Medical Informatics Assoc.4
2014 Decision support from local data: Creating adaptive order menus from past clinician behavior
Jeffrey G. Klann, Peter Szolovits, Stephen M. Downs, Gunther Schadow
J. Biomed. Informatics2
2012 Prognostic Physiology: Modeling Patient Severity in Intensive Care Units Using Radial Domain Folding
Rohit Joshi, Peter Szolovits
AMIA2
2012 Using UMLS for Word Sense Disambiguation in Clinical Notes
Anna Rumshisky, Rachel Chasin, Özlem Uzuner, Peter Szolovits
AMIA4
2012 MCORES: a system for noun phrase coreference resolution for clinical records
abstract
OBJECTIVE: Narratives of electronic medical records contain information that can be useful for clinical practice and multi-purpose research. This information needs to be put into a structured form before it can be used by automated systems. Coreference resolution is a step in the transformation of narratives into a structured form. METHODS: This study presents a medical coreference resolution system (MCORES) for noun phrases in four frequently used clinical semantic categories: persons, problems, treatments, and tests. MCORES treats coreference resolution as a binary classification task. Given a pair of concepts from a semantic category, it determines coreferent pairs and clusters them into chains. MCORES uses an enhanced set of lexical, syntactic, and semantic features. Some MCORES features measure the distance between various representations of the concepts in a pair and can be asymmetric. RESULTS AND CONCLUSION: MCORES was compared with an in-house baseline that uses only single-perspective 'token overlap' and 'number agreement' features. MCORES was shown to outperform the baseline; its enhanced features contribute significantly to performance. In addition to the baseline, MCORES was compared against two available third-party, open-domain systems, RECONCILE(ACL09) and the Beautiful Anaphora Resolution Toolkit (BART). MCORES was shown to outperform both of these systems on clinical records.
Andreea Bodnari, Peter Szolovits, Özlem Uzuner
J. Am. Medical Informatics Assoc.2
2011 Marco Ramoni: an appreciation of academic achievement
abstract
We review the scholarly career of our colleague, Marco Ramoni, who died unexpectedly in the summer of 2010. His work mainly explored the development and application of Bayesian techniques to model clinical, public health, and bioinformatics questions. His contributions have led to improvements in our ability to model behavior that evolves in time, to explore systematic relationships among large sets of covariates, and to tease out the meaning of data on the role of genetic variation in the genesis of important diseases.
Isaac S. Kohane, Peter Szolovits
J. Am. Medical Informatics Assoc.2
2011 Care transitions as opportunities for clinicians to use data exchange services: how often do they occur?
abstract
BACKGROUND: The electronic exchange of health information among healthcare providers has the potential to produce enormous clinical benefits and financial savings, although realizing that potential will be challenging. The American Recovery and Reinvestment Act of 2009 will reward providers for 'meaningful use' of electronic health records, including participation in clinical data exchange, but the best ways to do so remain uncertain. METHODS: We analyzed patient visits in one community in which a high proportion of providers were using an electronic health record and participating in data exchange. Using claims data from one large private payer for individuals under age 65 years, we computed the number of visits to a provider which involved transitions in care from other providers as a percentage of total visits. We calculated this 'transition percentage' for individual providers and medical groups. RESULTS: On average, excluding radiology and pathology, approximately 51% of visits involved care transitions between individual providers in the community and 36%-41% involved transitions between medical groups. There was substantial variation in transition percentage across medical specialties, within specialties and across medical groups. Specialists tended to have higher transition percentages and smaller ranges within specialty than primary care physicians, who ranged from 32% to 95% (including transitions involving radiology and pathology). The transition percentages of pediatric practices were similar to those of adult primary care, except that many transitions occurred among pediatric physicians within a single medical group. CONCLUSIONS: Care transition patterns differed substantially by type of practice and should be considered in designing incentives to foster providers' meaningful use of health data exchange services.
Robert S. Rudin, Claudia A. Salzberg, Peter Szolovits, Lynn A. Volk, Steven R. Simon, David W. Bates
J. Am. Medical Informatics Assoc.3
2011 Possibilities for Healthcare Computing
Peter Szolovits
J. Comput. Sci. Technol.1
2009 ICU Acuity: Real-time Models versus Daily Models
Caleb W. Hug, Peter Szolovits
AMIA2
2009 The coming of age of artificial intelligence in medicine
Vimla L. Patel, Edward H. Shortliffe, Mario Stefanelli, Peter Szolovits, Michael R. Berthold, Riccardo Bellazzi, Ameen Abu-Hanna
Artif. Intell. Medicine4
2008 A de-identifier for medical discharge summaries
Özlem Uzuner, Tawanda C. Sibanda, Yuan Luo 0001, Peter Szolovits
Artif. Intell. Medicine4
2008 Patient-specific learning in real time for adaptive monitoring in critical care
Peter Szolovits
J. Biomed. Informatics2
2007 Comment: What Is a Grid?
abstract
Precision of language is often thought to contribute to precision of thought, and certainly helps to communicate ideas unambiguously. A countervailing tendency, however, causes people to adopt terms developed in one field to stand for analogous concepts that may relate only incidentally to the original. Eventually the meaning of the original term becomes so broad and heterogeneous that we recognize its imprecise use as an impediment to communication, and the community adopts more precise language to clarify meaning. I argue that it is time to apply this corrective process to the term “grid.” This suggestion arises from my own confusion listening to numerous talks at the 2006 AMIA Symposium (and elsewhere), where speakers describe grids that have little in common. The Compact Oxford English Dictionary defines “grid” with four noun meanings,1 all derived from “gridiron,” a griddle for grilling meat: a framework of spaced bars that are parallel to or cross each other. a network of lines that cross each other to form a series of squares or rectangles. a network of cables or pipes for distributing power, especially high-voltage electricity. a pattern of lines marking the starting places on a motor-racing track. The term “grid computing” was adopted in the 1990's to describe an architecture and set of communication and policy standards to allow large groups of “personal” computers to work together to provide at low cost the computational power of a supercomputer by exploiting parallelism. This approach was pioneered by the scientific computing community, and has been formalized through the efforts of the Globus Alliance.2 This meaning of “grid” is now broadly accepted and, except for some technical variations, is used fairly unambiguously. At AMIA and in various other venues, however, I hear “grid” used to mean a very broad range of goals and methods: a community of common interests, a social and funding infrastructure to encourage data sharing, a technical approach to what we used to call federated databases, standardization and ontology construction for specific fields, and of course “real” grid computing, in its original meaning. For example, most descriptions of caBIG, including plenary talks at AMIA, present that project as addressing the first two of the above meanings: “caBIG™ is a voluntary network or grid connecting individuals and institutions to enable the sharing of data and tools, creating a World Wide Web of cancer research.”3 The Medical Grid project seems to take a narrower view, aiming to provide a set of federated tools for imaging and analysis.4 Yet others seek an integration of evolving ideas of ontology construction, the semantic Web, service oriented architectures, and grid computing.5 And of course many projects exploit formal grid architectures to achieve large-scale computing.6 I will be happy to leave to others in the community the crystallization of exactly which topics deserve a new name and structure, but I believe that we should enhance the clarity of our discussions by finding distinct names for distinct ideas. Peter Szolovits MIT
Peter Szolovits
J. Am. Medical Informatics Assoc.1
2007 Viewpoint Paper: Evaluating the State-of-the-Art in Automatic De-identification
abstract
To facilitate and survey studies in automatic de-identification, as a part of the i2b2 (Informatics for Integrating Biology to the Bedside) project, authors organized a Natural Language Processing (NLP) challenge on automatically removing private health information (PHI) from medical discharge records. This manuscript provides an overview of this de-identification challenge, describes the data and the annotation process, explains the evaluation metrics, discusses the nature of the systems that addressed the challenge, analyzes the results of received system runs, and identifies directions for future research. The de-indentification challenge data consisted of discharge summaries drawn from the Partners Healthcare system. Authors prepared this data for the challenge by replacing authentic PHI with synthesized surrogates. To focus the challenge on non-dictionary-based de-identification methods, the data was enriched with out-of-vocabulary PHI surrogates, i.e., made up names. The data also included some PHI surrogates that were ambiguous with medical non-PHI terms. A total of seven teams participated in the challenge. Each team submitted up to three system runs, for a total of sixteen submissions. The authors used precision, recall, and F-measure to evaluate the submitted system runs based on their token-level and instance-level performance on the ground truth. The systems with the best performance scored above 98% in F-measure for all categories of PHI. Most out-of-vocabulary PHI could be identified accurately. However, identifying ambiguous PHI proved challenging. The performance of systems on the test data set is encouraging. Future evaluations of these systems will involve larger data sets from more heterogeneous sources.
Özlem Uzuner, Yuan Luo 0001, Peter Szolovits
J. Am. Medical Informatics Assoc.3
2006 Syntactically-Informed Semantic Category Recognizer for Discharge Summaries
Tawanda C. Sibanda, Peter Szolovits, Özlem Uzuner
AMIA3
2005 Copy Fees and Patients' Rights to Obtain a Copy of Their Medical Records: From Law to Reality
Gianluigi Fioriglio, Peter Szolovits
AMIA2
2003 Adding a Medical Lexicon to an English Parser
Peter Szolovits
AMIA1
2001 Informatics Support for the Management of Drug Resistant Tuberculosis in Peru and Russia
Hamish S. F. Fraser, Libby Levison, Michael Nikiforov, Darius Jazayeri, Candy Day, Peter Szolovits, Jim Y. Kim
AMIA6
2001 Secure Health Information Sharing System: SHARE
Lik Mui, Mojdeh Mohtashemi, Peter Szolovits
AMIA4
2000 Appropriate Technology for Telemedicine in Developing Countries
Hamish S. F. Fraser, Darius Jazayeri, Peter Szolovits, St John D. McGrath
AMIA3
2000 First Steps Towards Implementing an International Training Program in Medical Informatics: The Brazil/USA Project
Lucila Ohno-Machado, Aziz A. Boxwala, Hamish S. F. Fraser, Robert A. Greenes, Isaac S. Kohane, Heimar F. Marin, Eduardo P. Marques, Eduardo Massad, Beatriz H. S. C. Rocha, Roberto A. Rocha, Laura M. Smeaton, Peter Szolovits
AMIA12
1998 Health information identification and de-identification toolkit
Isaac S. Kohane, Hongmei Dong, Peter Szolovits
AMIA3
1997 A Visual Method for Input of Uncertain Time-Oriented Data
Peter Szolovits
AMIA2
1997 Confidentiality of Medical Records in the W3-EMRS Project
David M. Rind, Isaac S. Kohane, Peter Szolovits, Charles Safran, Henry C. Chueh, G. Octo Barnett
AMIA3
1997 A Java-based multi-institutional medical information retrieval system
F. J. van Wingerde, Karen L. Bradshaw, Peter Szolovits, Isaac S. Kohane
AMIA4
1997 Application of Information Technology: A WWW Implementation of National Recommendations for Protecting Electronic Health Information
abstract
In March of 1997, the National Research Council (NRC) of the National Academy of Sciences issued the report, "For the Record: Protecting Electronic Health Information." Concluding that the current practices at the majority of health care facilities in the United States are insufficient, the Council delineated both technical and organizational approaches to protecting electronic health information. The Beth Israel Deaconess Medical Center recently implemented a proof-of-concept, Web-based, cross-institutional medical record, CareWeb, which incorporates the NRC security and confidentiality recommendations. We report on our WWW implementation of the NRC recommendations and an initial evaluation of the balance between ease of use and confidentiality.
John D. Halamka, Peter Szolovits, David M. Rind, Charles Safran
J. Am. Medical Informatics Assoc.2
1996 Application of Technology: Building National Electronic Medical Record Systems via the World Wide Web
abstract
Electronic medical record systems (EMRSs) currently do not lend themselves easily to cross-institutional clinical care and research. Unique system designs coupled with a lack of standards have led to this difficulty. The authors have designed a preliminary EMRS architecture (W3-EMRS) that exploits the multiplatform, multiprotocol, client-server technology of the World Wide Web. The architecture abstracts the clinical information model and the visual presentation away from the underlying EMRS. As a result, computation upon data elements of the EMRS and their presentation are no longer tied to the underlying EMRS structures. The architecture is intended to enable implementation of programs that provide uniform access to multiple, heterogeneous legacy EMRSs. The authors have implemented an initial prototype of W3-EMRS that accesses the database of the Boston Children's Hospital Clinician's Workstation.
Isaac S. Kohane, Philip Greenspun, James C. Fackler, Christopher Cimino, Peter Szolovits
J. Am. Medical Informatics Assoc.5
1994 Global Conditioning for Probabilistic Inference in Belief Networks
Ross D. Shachter, Stig K. Andersen, Peter Szolovits
UAI3
1994 Policy Forum: Against Simple Universal Health-Care Identifiers
abstract
Peter Szolovits, PhD, Isaac Kohane, MD, PhD; Against Simple Universal Health-care Identifiers, Journal of the American Medical Informatics Association, Volume 1
Peter Szolovits, Isaac S. Kohane
J. Am. Medical Informatics Assoc.1
1993 I-in-a-Box: A Knowledge-Based System for Space Science Experimentation
Richard Frainier, Nicolas Groleau, Lyman Hazelton, Peter Szolovits, Laurence Young, Silvano Colombano, Irving C. Statler, Michael Compton II
IAAI4
1993 Categorical and Probabilistic Reasoning in Medicine Revisited
Peter Szolovits, Stephen G. Pauker
Artif. Intell.1
1982 Information Acquisition in Diagnosis
Ramesh S. Patil, Peter Szolovits, William B. Schwartz
AAAI2
1981 Causal Understanding of Patient Illness in Medical Diagnosis
Ramesh S. Patil, Peter Szolovits, William B. Schwartz
IJCAI2
1981 Brand X: LISP Suport for Semantic Networks
Peter Szolovits, William A. Martin
IJCAI1
1978 Categorical and Probabilistic Reasoning in Medical Diagnosis
Peter Szolovits, Stephen G. Pauker
Artif. Intell.1
1974 The REL animated film language
abstract
Motion picture films generated by computers have now been produced by a number of groups across the country. We wish to report here on the design and implementation of the REL Animated Film Language (AFL), designed primarily for use by artists interested in the aesthetics of abstract motion graphics. The language is simple enough to be used by the artist directly, rather than by a professional programmer acting for the artist. It provides convenient means of expressing spatial and temporal changes in the shape and location of the objects with which the artist constructs his compositions. The artist uses the language to express the inter-object relationships which embody the aesthetic content of the work.
Frederick B. Thompson, Richard H. Bigelow, Norton Greenfeld, J. Odden, D. Reece, Peter Szolovits
SIGGRAPH6