VLDB 2026 Research / reviewers in the wild / expert
Guergana K. Savova
dblp:73/5451 · also Guergana Savova
· DBLP profile ↗
69ranked-venue papers
14as first author
7since 2021 · last 2026
0000-0002-5887-200XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 58 · 12 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dataset of Psychiatric Hospital Notes with Temporal Information AnnotationsabstractTemporal information extraction is the task of identifying temporal entities in a text and relating them to each other. In medicine, electronic health records (EHRs) contain text that documents the sequence of events during an encounter with a patient, and sometimes the events prior to the encounter (e.g., psychosocial environment and history). Temporality is especially important for the specialty of psychiatry. In this work, we describe the updates to the guidelines that allowed us to create a corpus of temporally-annotated psychiatric discharge summaries and progress notes in English. These updated guidelines were used to create a corpus of over 18,000 events, 2,200 time expressions, and 13,000 temporal relations. Temporal information extraction performance with a baseline system trained on non-psychiatric data obtains an F1 score of 0.152 on relation extraction, indicating the importance of this new dataset for making progress on temporal information extraction in the psychiatric domain. Timothy A. Miller, Gaby Dinh, David Harris 0004, Wonjin Yoon, Spencer Thomas, Boyu Ren, Mei-Hua Hall, Guergana K. Savova |
LREC | 8 |
| 2025 | Aspect-Oriented Summarization for Psychiatric Short-Term Readmission PredictionabstractRecent progress in large language models (LLMs) has enabled the automated processing of lengthy documents even without supervised training on a task-specific dataset.Yet, their zero-shot performance in complex tasks as opposed to straightforward information extraction tasks remains suboptimal.One feasible approach for tasks with lengthy, complex input is to first summarize the document and then apply supervised fine-tuning to the summary.However, the summarization process inevitably results in some loss of information.In this study we present a method for processing the summaries of long documents aimed to capture different important aspects of the original document.We hypothesize that LLM summaries generated with different aspect-oriented prompts contain different information signals, and we propose methods to measure these differences.We introduce approaches to effectively integrate signals from these different summaries for supervised training of transformer models.We validate our hypotheses on a high-impact task -30-day readmission prediction from a psychiatric discharge -using real-world data from four hospitals, and show that our proposed method increases the prediction performance for the complex task of predicting patient outcome. Wonjin Yoon, Boyu Ren, Spencer Thomas, Chanhwi Kim, Guergana K. Savova, Mei-Hua Hall, Timothy A. Miller |
EMNLP | 5 |
| 2025 | Identifying task groupings for multi-task learning using pointwise V-usable information
Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
J. Biomed. Informatics | 4 |
| 2024 | Evaluating the ChatGPT family of models for biomedical reasoning and classificationabstractOBJECTIVE: Large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates ChatGPT family of models (GPT-3.5, GPT-4) in biomedical tasks beyond question-answering. MATERIALS AND METHODS: We evaluated model performance with 11 122 samples for two fundamental tasks in the biomedical domain-classification (n = 8676) and reasoning (n = 2446). The first task involves classifying health advice in scientific literature, while the second task is detecting causal relations in biomedical literature. We used 20% of the dataset for prompt development, including zero- and few-shot settings with and without chain-of-thought (CoT). We then evaluated the best prompts from each setting on the remaining dataset, comparing them to models using simple features (BoW with logistic regression) and fine-tuned BioBERT models. RESULTS: Fine-tuning BioBERT produced the best classification (F1: 0.800-0.902) and reasoning (F1: 0.851) results. Among LLM approaches, few-shot CoT achieved the best classification (F1: 0.671-0.770) and reasoning (F1: 0.682) results, comparable to the BoW model (F1: 0.602-0.753 and 0.675 for classification and reasoning, respectively). It took 78 h to obtain the best LLM results, compared to 0.078 and 0.008 h for the top-performing BioBERT and BoW models, respectively. DISCUSSION: The simple BoW model performed similarly to the most complex LLM prompting. Prompt engineering required significant investment. CONCLUSION: Despite the excitement around viral ChatGPT, fine-tuning for two fundamental biomedical natural language processing tasks remained the best strategy. Shan Chen 0004, Yingya Li, Hoang Van, Hugo J. W. L. Aerts, Guergana K. Savova, Danielle S. Bitterman |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language ModelsabstractThe bias-variance tradeoff is the idea that learning methods need to balance model complexity with data size to minimize both under-fitting and over-fitting.Recent empirical work and theoretical analyses with over-parameterized neural networks challenge the classic bias-variance trade-off notion suggesting that no such trade-off holds: as the width of the network grows, bias monotonically decreases while variance initially increases followed by a decrease.In this work, we first provide a variance decomposition-based justification criteria to examine whether large pretrained neural models in a fine-tuning setting are generalizable enough to have low bias and variance.We then perform theoretical and empirical analysis using ensemble methods explicitly designed to decrease variance due to optimization.This results in essentially a two-stage fine-tuning algorithm that first ratchets down bias and variance iteratively, and then uses a selected fixed-bias model to further reduce variance due to optimization by ensembling.We also analyze the nature of variance change with the ensemble size in low-and high-resource classes.Empirical results show that this two-stage method obtains strong results on SuperGLUE tasks and clinical information extraction tasks.Code and settings are available: https://github.com/christa60/ bias-var-fine-tuning-plms.git Lijing Wang 0001, Yingya Li, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
ACL (1) | 5 |
| 2023 | Improving model transferability for clinical note section classification models using continued pretrainingabstractOBJECTIVE: The classification of clinical note sections is a critical step before doing more fine-grained natural language processing tasks such as social determinants of health extraction and temporal information extraction. Often, clinical note section classification models that achieve high accuracy for 1 institution experience a large drop of accuracy when transferred to another institution. The objective of this study is to develop methods that classify clinical note sections under the SOAP ("Subjective," "Object," "Assessment," and "Plan") framework with improved transferability. MATERIALS AND METHODS: We trained the baseline models by fine-tuning BERT-based models, and enhanced their transferability with continued pretraining, including domain-adaptive pretraining and task-adaptive pretraining. We added in-domain annotated samples during fine-tuning and observed model performance over a varying number of annotated sample size. Finally, we quantified the impact of continued pretraining in equivalence of the number of in-domain annotated samples added. RESULTS: We found continued pretraining improved models only when combined with in-domain annotated samples, improving the F1 score from 0.756 to 0.808, averaged across 3 datasets. This improvement was equivalent to adding 35 in-domain annotated samples. DISCUSSION: Although considered a straightforward task when performing in-domain, section classification is still a considerably difficult task when performing cross-domain, even using highly sophisticated neural network-based methods. CONCLUSION: Continued pretraining improved model transferability for cross-domain clinical note section classification in the presence of a small amount of in-domain labeled samples. Weipeng Zhou, Meliha Yetisgen, Majid Afshar, Yanjun Gao, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | DeepPhe: Natural Language Processing Tools for Cancer Research and Surveillance
Harry Hochheiser, Sean Finan, Zhou Yuan, John D. Levander, Eric B. Durbin, Isaac Hands, Ramakanth Kavuluru, Jeremy L. Warner, Guergana K. Savova |
AMIA | 9 |
| 2020 | Training Women for Leadership in Informatics and Digital Health: A Report from the Inaugural Women in AMIA Leadership Program
Wendy W. Chapman, María Adela Grando, Merida L. Johns, Guergana K. Savova, Maia Hightower |
AMIA | 4 |
| 2020 | De-Identification of Clinical Text: Stakeholders' Perspectives and Acceptance of Automatic De-Identification
Stéphane M. Meystre, Jonathan C. Silverstein, Guergana K. Savova, Valentina Petkov, Bradley A. Malin |
AMIA | 3 |
| 2020 | Adverse drug event rates in pediatric pulmonary hypertension: a comparison of real-world data sourcesabstractOBJECTIVE: Real-world data (RWD) are increasingly used for pharmacoepidemiology and regulatory innovation. Our objective was to compare adverse drug event (ADE) rates determined from two RWD sources, electronic health records and administrative claims data, among children treated with drugs for pulmonary hypertension. MATERIALS AND METHODS: Textual mentions of medications and signs/symptoms that may represent ADEs were identified in clinical notes using natural language processing. Diagnostic codes for the same signs/symptoms were identified in our electronic data warehouse for the patients with textual evidence of taking pulmonary hypertension-targeted drugs. We compared rates of ADEs identified in clinical notes to those identified from diagnostic code data. In addition, we compared putative ADE rates from clinical notes to those from a healthcare claims dataset from a large, national insurer. RESULTS: Analysis of clinical notes identified up to 7-fold higher ADE rates than those ascertained from diagnostic codes. However, certain ADEs (eg, hearing loss) were more often identified in diagnostic code data. Similar results were found when ADE rates ascertained from clinical notes and national claims data were compared. DISCUSSION: While administrative claims and clinical notes are both increasingly used for RWD-based pharmacovigilance, ADE rates substantially differ depending on data source. CONCLUSION: Pharmacovigilance based on RWD may lead to discrepant results depending on the data source analyzed. Further work is needed to confirm the validity of identified ADEs, to distinguish them from disease effects, and to understand tradeoffs in sensitivity and specificity between data sources. Alon Geva, Steven H. Abman, Shannon F. Manzi, Dunbar D. Ivy, Mary Mullen, John Griffin, Chen Lin 0002, Guergana K. Savova, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 8 |
| 2020 | Does BERT need domain adaptation for clinical negation detection?abstractINTRODUCTION: Classifying whether concepts in an unstructured clinical text are negated is an important unsolved task. New domain adaptation and transfer learning methods can potentially address this issue. OBJECTIVE: We examine neural unsupervised domain adaptation methods, introducing a novel combination of domain adaptation with transformer-based transfer learning methods to improve negation detection. We also want to better understand the interaction between the widely used bidirectional encoder representations from transformers (BERT) system and domain adaptation methods. MATERIALS AND METHODS: We use 4 clinical text datasets that are annotated with negation status. We evaluate a neural unsupervised domain adaptation algorithm and BERT, a transformer-based model that is pretrained on massive general text datasets. We develop an extension to BERT that uses domain adversarial training, a neural domain adaptation method that adds an objective to the negation task, that the classifier should not be able to distinguish between instances from 2 different domains. RESULTS: The domain adaptation methods we describe show positive results, but, on average, the best performance is obtained by plain BERT (without the extension). We provide evidence that the gains from BERT are likely not additive with the gains from domain adaptation. DISCUSSION: Our results suggest that, at least for the task of clinical negation detection, BERT subsumes domain adaptation, implying that BERT is already learning very general representations of negation phenomena such that fine-tuning even on a specific corpus does not lead to much overfitting. CONCLUSION: Despite being trained on nonclinical text, the large training sets of models like BERT lead to large gains in performance for the clinical negation detection task. Chen Lin 0002, Steven Bethard, Dmitriy Dligach, Farig Sadeque, Guergana K. Savova, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Supervised methods to extract clinical events from cardiology reports in Italian
Natalia Viani, Timothy A. Miller, Carlo Napolitano, Silvia G. Priori, Guergana K. Savova, Riccardo Bellazzi, Lucia Sacchi |
J. Biomed. Informatics | 5 |
| 2018 | Toward Large-scale and Multi-facet Analysis of First Person Alcohol Drinking
Hadi Amiri, Kara M. Magane, Lauren E. Wisk, Guergana K. Savova, Elissa R. Weitzman |
AMIA | 4 |
| 2018 | ADEPT: An End-to-End System for High-Throughput Pharmacovigilance
Alon Geva, Jason Stedman, Guergana K. Savova, Chen Lin 0002, Shannon F. Manzi, Paul Avillach, Kenneth D. Mandl |
AMIA | 3 |
| 2018 | Women in AMIA - Resources for Emerging Leaders
Guergana K. Savova, Merida L. Johns, Nancy M. Lorenzi, Patricia Flatley Brennan, Rebecca S. Jacobson |
AMIA | 1 |
| 2018 | Computable Longitudinal Patient Trajectories
Jeremy L. Warner, Guergana K. Savova, Noémie Elhadad, Lisa Bastarache, David Gotz |
AMIA | 2 |
| 2018 | Spotting Spurious Data with Neural NetworksabstractHadi Amiri, Timothy Miller, Guergana Savova. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hadi Amiri, Timothy A. Miller, Guergana K. Savova |
NAACL-HLT | 3 |
| 2017 | Recurrent Neural Network Architectures for Event Extraction from Italian Medical Reports
Natalia Viani, Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Carlo Napolitano, Silvia G. Priori, Riccardo Bellazzi, Lucia Sacchi, Guergana K. Savova |
AIME | 9 |
| 2017 | Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Noémie Elhadad, Sumithra Velupillai, Hua Xu 0001, Guergana K. Savova |
AMIA | 5 |
| 2017 | DeepPhe - A Natural Language Processing System for Extracting Cancer Phenotypes from Clinical Records
Guergana K. Savova, Eugene Tseytlin, Sean Finan, Melissa Castine, Timothy A. Miller, Olga Medvedeva, David Harris 0004, Harry Hochheiser, Chen Lin 0002, Girish Chavan, Rebecca S. Jacobson |
AMIA | 1 |
| 2017 | Repeat before Forgetting: Spaced Repetition for Efficient and Effective Training of Neural NetworksabstractWe present a novel approach for training artificial neural networks.Our approach is inspired by broad evidence in psychology that shows human learners can learn efficiently and effectively by increasing intervals of time between subsequent reviews of previously learned materials (spaced repetition).We investigate the analogy between training neural models and findings in psychology about human memory model and develop an efficient and effective algorithm to train neural models.The core part of our algorithm is a cognitively-motivated scheduler according to which training instances and their "reviews" are spaced over time.Our algorithm uses only 34-50% of data per epoch, is 2.9-4.8 times faster than standard training, and outperforms competing state-of-the-art baselines.1 Hadi Amiri, Timothy A. Miller, Guergana K. Savova |
EMNLP | 3 |
| 2017 | Towards generalizable entity-centric clinical coreference resolution
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Chen Lin 0002, Guergana K. Savova |
J. Biomed. Informatics | 5 |
| 2016 | Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001 |
AMIA | 7 |
| 2016 | Feature Portability in Cross-domain Clinical Coreference
Timothy A. Miller, Dmitriy Dligach, Chen Lin 0002, Steven Bethard, Guergana K. Savova |
AMIA | 5 |
| 2016 | PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportabilityabstractOBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data. Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 14 |
| 2016 | Multilayered temporal modeling for the clinical domainabstractOBJECTIVE: To develop an open-source temporal relation discovery system for the clinical domain. The system is capable of automatically inferring temporal relations between events and time expressions using a multilayered modeling strategy. It can operate at different levels of granularity--from rough temporality expressed as event relations to the document creation time (DCT) to temporal containment to fine-grained classic Allen-style relations. MATERIALS AND METHODS: We evaluated our systems on 2 clinical corpora. One is a subset of the Temporal Histories of Your Medical Events (THYME) corpus, which was used in SemEval 2015 Task 6: Clinical TempEval. The other is the 2012 Informatics for Integrating Biology and the Bedside (i2b2) challenge corpus. We designed multiple supervised machine learning models to compute the DCT relation and within-sentence temporal relations. For the i2b2 data, we also developed models and rule-based methods to recognize cross-sentence temporal relations. We used the official evaluation scripts of both challenges to make our results comparable with results of other participating systems. In addition, we conducted a feature ablation study to find out the contribution of various features to the system's performance. RESULTS: Our system achieved state-of-the-art performance on the Clinical TempEval corpus and was on par with the best systems on the i2b2 2012 corpus. Particularly, on the Clinical TempEval corpus, our system established a new F1 score benchmark, statistically significant as compared to the baseline and the best participating system. CONCLUSION: Presented here is the first open-source clinical temporal relation discovery system. It was built using a multilayered temporal modeling strategy and achieved top performance in 2 major shared tasks. Chen Lin 0002, Dmitriy Dligach, Timothy A. Miller, Steven Bethard, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Semi-supervised Learning for Phenotyping Tasks
Dmitriy Dligach, Timothy A. Miller, Guergana K. Savova |
AMIA | 3 |
| 2015 | Demonstrating the Advantages of Applying Data Mining Techniques on Time-Dependent Electronic Medical Records
Uri Kartoun, Vishesh Kumar, Su-Chun Cheng, Sheng Yu 0002, Katherine P. Liao, Elizabeth W. Karlson, Ashwin N. Ananthakrishnan, Zongqi Xia, Vivian S. Gainer, Andrew Cagan, Guergana K. Savova, Pei J. Chen, Shawn N. Murphy, Susanne E. Churchill, Isaac S. Kohane, Peter Szolovits, Tianxi Cai, Stanley Y. Shaw |
AMIA | 11 |
| 2015 | Robust Sentence Segmentation for Clinical Text
Timothy A. Miller, Sean Finan, Dmitriy Dligach, Guergana K. Savova |
AMIA | 4 |
| 2015 | Natural Language Processing for Phenotype Extraction: Challenges in Extraction and Representation
Guergana K. Savova, Rebecca S. Jacobson, Joshua C. Denny, Nicole L. Washington, Harry Hochheiser |
AMIA | 1 |
| 2015 | Automatic identification of methotrexate-induced liver toxicity in patients with rheumatoid arthritis from the electronic medical recordabstractOBJECTIVES: To improve the accuracy of mining structured and unstructured components of the electronic medical record (EMR) by adding temporal features to automatically identify patients with rheumatoid arthritis (RA) with methotrexate-induced liver transaminase abnormalities. MATERIALS AND METHODS: Codified information and a string-matching algorithm were applied to a RA cohort of 5903 patients from Partners HealthCare to select 1130 patients with potential liver toxicity. Supervised machine learning was applied as our key method. For features, Apache clinical Text Analysis and Knowledge Extraction System (cTAKES) was used to extract standard vocabulary from relevant sections of the unstructured clinical narrative. Temporal features were further extracted to assess the temporal relevance of event mentions with regard to the date of transaminase abnormality. All features were encapsulated in a 3-month-long episode for classification. Results were summarized at patient level in a training set (N=480 patients) and evaluated against a test set (N=120 patients). RESULTS: The system achieved positive predictive value (PPV) 0.756, sensitivity 0.919, F1 score 0.829 on the test set, which was significantly better than the best baseline system (PPV 0.590, sensitivity 0.703, F1 score 0.642). Our innovations, which included framing the phenotype problem as an episode-level classification task, and adding temporal information, all proved highly effective. CONCLUSIONS: Automated methotrexate-induced liver toxicity phenotype discovery for patients with RA based on structured and unstructured information in the EMR shows accurate results. Our work demonstrates that adding temporal features significantly improved classification results. Chen Lin 0002, Elizabeth W. Karlson, Dmitriy Dligach, Monica P. Ramirez, Timothy A. Miller, Huan Mo, Natalie S. Braggs, Andrew Cagan, Vivian S. Gainer, Joshua C. Denny, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 11 |
| 2015 | Evaluating the state of the art in disorder recognition and normalization of the clinical narrativeabstractOBJECTIVE: The ShARe/CLEF eHealth 2013 Evaluation Lab Task 1 was organized to evaluate the state of the art on the clinical text in (i) disorder mention identification/recognition based on Unified Medical Language System (UMLS) definition (Task 1a) and (ii) disorder mention normalization to an ontology (Task 1b). Such a community evaluation has not been previously executed. Task 1a included a total of 22 system submissions, and Task 1b included 17. Most of the systems employed a combination of rules and machine learners. MATERIALS AND METHODS: We used a subset of the Shared Annotated Resources (ShARe) corpus of annotated clinical text--199 clinical notes for training and 99 for testing (roughly 180 K words in total). We provided the community with the annotated gold standard training documents to build systems to identify and normalize disorder mentions. The systems were tested on a held-out gold standard test set to measure their performance. RESULTS: For Task 1a, the best-performing system achieved an F1 score of 0.75 (0.80 precision; 0.71 recall). For Task 1b, another system performed best with an accuracy of 0.59. DISCUSSION: Most of the participating systems used a hybrid approach by supplementing machine-learning algorithms with features generated by rules and gazetteers created from the training data and from external resources. CONCLUSIONS: The task of disorder normalization is more challenging than that of identification. The ShARe corpus is available to the community as a reference standard for future studies. Sameer Pradhan, Noémie Elhadad, Brett R. South, David Martínez 0001, Lee M. Christensen, Amy Vogel, Hanna Suominen, Wendy W. Chapman, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 9 |
| 2014 | Narrative Event and Temporal Relation Visualization Tool
Sean Finan, Piet C. de Groen, Guergana K. Savova |
AMIA | 3 |
| 2014 | Developing a Section Labeler for Clinical Documents
Peter J. Haug, Xinzi Wu, Jeffrey P. Ferraro, Guergana K. Savova, Stanley M. Huff, Christopher G. Chute |
AMIA | 4 |
| 2014 | Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Hercules Dalianis, Guergana K. Savova, Pierre Zweigenbaum |
AMIA | 3 |
| 2014 | Disease/Disorder Semantic Template Filling - Information Extraction Challenge in the ShARe/CLEF eHealth Evaluation Lab 2014
Sumithra Velupillai, Danielle L. Mowery, Lee M. Christensen, Noémie Elhadad, Sameer Pradhan, Guergana K. Savova, Wendy W. Chapman |
AMIA | 6 |
| 2014 | Discovering body site and severity modifiers in clinical textsabstractOBJECTIVE: To research computational methods for discovering body site and severity modifiers in clinical texts. METHODS: We cast the task of discovering body site and severity modifiers as a relation extraction problem in the context of a supervised machine learning framework. We utilize rich linguistic features to represent the pairs of relation arguments and delegate the decision about the nature of the relationship between them to a support vector machine model. We evaluate our models using two corpora that annotate body site and severity modifiers. We also compare the model performance to a number of rule-based baselines. We conduct cross-domain portability experiments. In addition, we carry out feature ablation experiments to determine the contribution of various feature groups. Finally, we perform error analysis and report the sources of errors. RESULTS: The performance of our method for discovering body site modifiers achieves F1 of 0.740-0.908 and our method for discovering severity modifiers achieves F1 of 0.905-0.929. DISCUSSION: Results indicate that both methods perform well on both in-domain and out-domain data, approaching the performance of human annotators. The most salient features are token and named entity features, although syntactic dependency features also contribute to the overall performance. The dominant sources of errors are infrequent patterns in the data and inability of the system to discern deeper semantic structures. CONCLUSIONS: We investigated computational methods for discovering body site and severity modifiers in clinical texts. Our best system is released open source as part of the clinical Text Analysis and Knowledge Extraction System (cTAKES). Dmitriy Dligach, Steven Bethard, Lee Becker, Timothy A. Miller, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 5 |
| 2014 | Temporal Annotation in the Clinical DomainabstractThis article discusses the requirements of a formal specification for the annotation of temporal information in clinical narratives. We discuss the implementation and extension of ISO-TimeML for annotating a corpus of clinical notes, known as the THYME corpus. To reflect the information task and the heavily inference-based reasoning demands in the domain, a new annotation guideline has been developed, "the THYME Guidelines to ISO-TimeML (THYME-TimeML)". To clarify what relations merit annotation, we distinguish between linguistically-derived and inferentially-derived temporal orderings in the text. We also apply a top performing TempEval 2013 system against this new resource to measure the difficulty of adapting systems to the clinical domain. The corpus is available to the community and has been proposed for use in a SemEval 2015 task. William F. Styler IV, Steven Bethard, Sean Finan, Martha Palmer, Sameer Pradhan, Piet C. de Groen, Bradley James Erickson, Timothy A. Miller, Chen Lin 0002, Guergana K. Savova, James Pustejovsky |
Trans. Assoc. Comput. Linguistics | 10 |
| 2013 | Discovering Body Site and Severity Modifiers in Clinical Texts
Dmitriy Dligach, Timothy A. Miller, Guergana K. Savova |
AMIA | 3 |
| 2013 | Automatic Prediction of Rheumatoid Arthritis Disease Activity from the Electronic Medical Records
Chen Lin 0002, Elizabeth W. Karlson, Helena Canhão, Timothy A. Miller, Dmitriy Dligach, Pei J. Chen, Raúl N. Pérez, Michael E. Weinblatt, Nancy A. Shadick, Robert M. Plenge, Guergana K. Savova |
AMIA | 12 |
| 2013 | Discovering Time Expressions in Clinical Text
Timothy A. Miller, Dmitriy Dligach, Steven Bethard, Sameer Pradhan, Chen Lin 0002, Guergana K. Savova |
AMIA | 6 |
| 2013 | Panel: Shared Resources, Shared Code, and Shared Activities in Clinical Natural Language Processing
Guergana K. Savova, Wendy W. Chapman, Noémie Elhadad, Martha Palmer |
AMIA | 1 |
| 2013 | Towards comprehensive syntactic and semantic annotations of the clinical narrativeabstractOBJECTIVE: To create annotated clinical narratives with layers of syntactic and semantic labels to facilitate advances in clinical natural language processing (NLP). To develop NLP algorithms and open source components. METHODS: Manual annotation of a clinical narrative corpus of 127 606 tokens following the Treebank schema for syntactic information, PropBank schema for predicate-argument structures, and the Unified Medical Language System (UMLS) schema for semantic information. NLP components were developed. RESULTS: The final corpus consists of 13 091 sentences containing 1772 distinct predicate lemmas. Of the 766 newly created PropBank frames, 74 are verbs. There are 28 539 named entity (NE) annotations spread over 15 UMLS semantic groups, one UMLS semantic type, and the Person semantic category. The most frequent annotations belong to the UMLS semantic groups of Procedures (15.71%), Disorders (14.74%), Concepts and Ideas (15.10%), Anatomy (12.80%), Chemicals and Drugs (7.49%), and the UMLS semantic type of Sign or Symptom (12.46%). Inter-annotator agreement results: Treebank (0.926), PropBank (0.891-0.931), NE (0.697-0.750). The part-of-speech tagger, constituency parser, dependency parser, and semantic role labeler are built from the corpus and released open source. A significant limitation uncovered by this project is the need for the NLP community to develop a widely agreed-upon schema for the annotation of clinical concepts and their relations. CONCLUSIONS: This project takes a foundational step towards bringing the field of clinical NLP up to par with NLP in the general domain. The corpus creation and NLP components provide a resource for research and application development that would have been previously impossible. Daniel Albright, Arrick Lanfranchi, Anwen Fredriksen, William F. Styler IV, Colin Warner, Jena D. Hwang, Jinho D. Choi, Dmitriy Dligach, Rodney D. Nielsen, James H. Martin, Wayne H. Ward, Martha Palmer, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 13 |
| 2013 | Large-scale evaluation of automated clinical note de-identification and its impact on information extractionabstractOBJECTIVE: (1) To evaluate a state-of-the-art natural language processing (NLP)-based approach to automatically de-identify a large set of diverse clinical notes. (2) To measure the impact of de-identification on the performance of information extraction algorithms on the de-identified documents. MATERIAL AND METHODS: A cross-sectional study that included 3503 stratified, randomly selected clinical notes (over 22 note types) from five million documents produced at one of the largest US pediatric hospitals. Sensitivity, precision, F value of two automated de-identification systems for removing all 18 HIPAA-defined protected health information elements were computed. Performance was assessed against a manually generated 'gold standard'. Statistical significance was tested. The automated de-identification performance was also compared with that of two humans on a 10% subsample of the gold standard. The effect of de-identification on the performance of subsequent medication extraction was measured. RESULTS: The gold standard included 30 815 protected health information elements and more than one million tokens. The most accurate NLP method had 91.92% sensitivity (R) and 95.08% precision (P) overall. The performance of the system was indistinguishable from that of human annotators (annotators' performance was 92.15%(R)/93.95%(P) and 94.55%(R)/88.45%(P) overall while the best system obtained 92.91%(R)/95.73%(P) on same text). The impact of automated de-identification was minimal on the utility of the narrative notes for subsequent information extraction as measured by the sensitivity and precision of medication name extraction. DISCUSSION AND CONCLUSION: NLP-based de-identification shows excellent performance that rivals the performance of human annotators. Furthermore, unlike manual de-identification, the automated approach scales up to millions of documents quickly and inexpensively. Louise Deléger, Katalin Molnár, Guergana K. Savova, Fei Xia 0004, Todd Lingren, Qi Li 0004, Keith Marsolo, Anil G. Jegga, Megan Kaiser, Laura Stoutenborough, Imre Solti |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Automated discovery of drug treatment patterns for endocrine therapy of breast cancer within an electronic medical recordabstractOBJECTIVE: To develop an algorithm for the discovery of drug treatment patterns for endocrine breast cancer therapy within an electronic medical record and to test the hypothesis that information extracted using it is comparable to the information found by traditional methods. MATERIALS: The electronic medical charts of 1507 patients diagnosed with histologically confirmed primary invasive breast cancer. METHODS: The automatic drug treatment classification tool consisted of components for: (1) extraction of drug treatment-relevant information from clinical narratives using natural language processing (clinical Text Analysis and Knowledge Extraction System); (2) extraction of drug treatment data from an electronic prescribing system; (3) merging information to create a patient treatment timeline; and (4) final classification logic. RESULTS: Agreement between results from the algorithm and from a nurse abstractor is measured for categories: (0) no tamoxifen or aromatase inhibitor (AI) treatment; (1) tamoxifen only; (2) AI only; (3) tamoxifen before AI; (4) AI before tamoxifen; (5) multiple AIs and tamoxifen cycles in no specific order; and (6) no specific treatment dates. Specificity (all categories): 96.14%-100%; sensitivity (categories (0)-(4)): 90.27%-99.83%; sensitivity (categories (5)-(6)): 0-23.53%; positive predictive values: 80%-97.38%; negative predictive values: 96.91%-99.93%. DISCUSSION: Our approach illustrates a secondary use of the electronic medical record. The main challenge is event temporality. CONCLUSION: We present an algorithm for automated treatment classification within an electronic medical record to combine information extracted through natural language processing with that extracted from structured databases. The algorithm has high specificity for all categories, high sensitivity for five categories, and low sensitivity for two categories. Guergana K. Savova, Janet E. Olson, Sean P. Murphy, Victoria L. Cafourek, Fergus J. Couch, Matthew P. Goetz, James N. Ingle, Vera J. Suman, Christopher G. Chute, Richard M. Weinshilboum |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | A system for coreference resolution for the clinical narrativeabstractOBJECTIVE: To research computational methods for coreference resolution in the clinical narrative and build a system implementing the best methods. METHODS: The Ontology Development and Information Extraction corpus annotated for coreference relations consists of 7214 coreferential markables, forming 5992 pairs and 1304 chains. We trained classifiers with semantic, syntactic, and surface features pruned by feature selection. For the three system components--for the resolution of relative pronouns, personal pronouns, and noun phrases--we experimented with support vector machines with linear and radial basis function (RBF) kernels, decision trees, and perceptrons. Evaluation of algorithms and varied feature sets was performed using standard metrics. RESULTS: The best performing combination is support vector machines with an RBF kernel and all features (MUC score=0.352, B(3)=0.690, CEAF=0.486, BLANC=0.596) outperforming a traditional decision tree baseline. DISCUSSION: The application showed good performance similar to performance on general English text. The main error source was sentence distances exceeding a window of 10 sentences between markables. A possible solution to this problem is hinted at by the fact that coreferent markables sometimes occurred in predictable (although distant) note sections. Another system limitation is failure to fully utilize synonymy and ontological knowledge. Future work will investigate additional ways to incorporate syntactic features into the coreference problem. CONCLUSION: We investigated computational methods for coreference resolution in the clinical narrative. The best methods are released as modules of the open source Clinical Text Analysis and Knowledge Extraction System and Ontology Development and Information Extraction platforms. Jiaping Zheng, Wendy W. Chapman, Timothy A. Miller, Chen Lin 0002, Rebecca S. Jacobson, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 6 |
| 2012 | Anaphoric reference in clinical reports: Characteristics of an annotated corpus
Wendy W. Chapman, Guergana K. Savova, Jiaping Zheng, Melissa Tharp, Rebecca S. Jacobson |
J. Biomed. Informatics | 2 |
| 2012 | Building a robust, scalable and standards-driven infrastructure for secondary use of EHR data: The SHARPn project
Susan Rea, Jyotishman Pathak, Guergana K. Savova, Thomas A. Oniki, Les Westberg, Calvin E. Beebe, Cui Tao, Craig G. Parker, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
J. Biomed. Informatics | 3 |
| 2011 | Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutionsabstractThis issue of JAMIA focuses on natural language processing (NLP) techniques for clinical-text information extraction. Several articles are offshoots of the yearly ‘Informatics for Integrating Biology and the Bedside’ (i2b2) (http://www.i2b2.org) NLP shared-task challenge, introduced by Uzuner et al (see page 552)1 and co-sponsored by the Veteran's Administration for the last 2 years. This shared task follows long-running challenge evaluations in other fields, such as the Message Understanding Conference (MUC) for information extraction,2 TREC3 for text information retrieval, and CASP4 for protein structure prediction. Shared tasks in the clinical domain are recent and include annual i2b2 Challenges that began in 2006, a challenge for multi-label classification of radiology reports sponsored by Cincinnati Children's Hospital in 2007,5 a 2011 Cincinnati Children's Hospital challenge on suicide notes,6 and the 2011 TREC information retrieval shared task involving retrieval of clinical cases from narrative records.7 Although NLP research in the clinical domain has been active since the 1960s, progress in the development of NLP applications for clinical text has been slow and lags behind progress in the general NLP domain. There are several barriers to NLP development in the clinical domain, and shared tasks like the i2b2/VA Challenge address some of these barriers. Nevertheless, many barriers remain and unless the community takes a more active role in developing novel approaches for addressing the barriers, advancement and innovation will continue to be slow. Historically, there have been substantial barriers to NLP development in the clinical domain. These barriers are not unique to the clinical domain: they also occur in the fields of software engineering and general NLP. Because of concerns regarding patient privacy and worry about revealing unfavorable institutional practices, hospitals and clinics have been extremely reluctant to allow access to clinical data for researchers from outside the associated institutions. The lack of reliable and inexpensive de-identification techniques for narrative reports has compounded the reluctance to share. Such restricted access to shared datasets has hindered collaboration and inhibited the ability to assess and adapt NLP technologies across institutions and among research groups. Several pioneering efforts5,8–11 have made clinical data available for sharing—we need more of these grass-roots efforts. Closely related but not completely conditional on lack of shared datasets is the deficiency of annotated clinical data for training NLP applications and benchmarking performance. The sublanguage of clinical reports often necessitates domain-specific development and training, and, as a consequence, NLP modules developed for general text typically do not perform as well on clinical narratives. We need increased coordination to create annotation sets that can be merged to produce larger training and evaluation sets. Without the ability to share data, the community has lacked incentives for developing common data models for manual and automatic annotations. The result is that annotated datasets are usually unique to the laboratory that generated them and thus remain small and that NLP modules that perform the same tasks cannot be substituted and compared without considerable translational effort. At present, the clinical NLP community is leveraging existing standards and conventions and working together to develop shared data models and to map annotations across information extraction applications. Adopting an existing NLP application or module is complicated—source code and documentation may be unavailable, and published descriptions may lack sufficient detail for reproducibility. Open source releases of clinical information extraction and retrieval systems have improved the opportunity to reproduce performance.12–15 Even with open source release, a tool may work less well in others' hands than in the hands of the original developers. Compounding the problem of reproducibility is the fact that proof-of-concept tools created in academic/research environments may not meet the highest software engineering quality, maintainability, scalability, or usability standards. And sometimes a tool may be over-fitted to a particular application, and modification to solve a similar problem may require wholesale changes. As Pedersen asserted,16 the NLP community needs to invest more in assisting others in applying and reproducing our results. In part due to previously listed barriers, collaboration within the clinical NLP community has been nominal. Development of NLP systems within the academic environment has centered around single institutions and single laboratories, and rather than building upon the foundations of previous work, the majority of clinical NLP systems developed over the last four decades have been reinvented as silos that are neither expanded nor applied outside of the individual laboratory. Other factors limiting collaboration include insufficient infrastructure for facilitating cooperation and the reality that collaboration is inherently inefficient. Nevertheless, as with the biomedical research community at large, a surge in progression beyond the last half century of research can only come through enhanced teamwork. A recent trend for teamwork across NLP research laboratories is evident in funded initiatives such as the VA's Consortium for Healthcare Informatics Research (CHIR)17 and in the ONC-funded SHARP Area 4 grant for Secondary Use of EHR data.18 Also, advances in open source development and conformity to common frameworks has led to recent advances in NLP allowing one team to extend work by another (eg, HiTEX12 built on GATE and cTAKES,14 ODIE,19 and Automated Retrieval Console (ARC) built on UIMA16). Although we are improving incrementally the predictive performance of clinical NLP tools, clinical NLP applications are seldom deployed in clinical, public health, or health services research settings. Currently, the perceived cost of applying NLP outweighs the perceived benefit. Deploying an NLP system typically requires a substantial amount of time from an expert NLP developer—normally, applications do not generalize and must be rebuilt, retrained, enhanced, and re-evaluated for each new task; the output of an NLP system typically requires extensive mapping to the specific problem being addressed; and the ability to aid a user in customizing the application is generally inadequate (see page 544).20 We need a shift of focus from accuracy in one task to generalizabiliy across many and from the production of papers as the sole output to production of usable software for medically relevant applications. We also need to understand where NLP tools fit into an overall user workflow so that the tools can be integrated into end-to-end applications for clinical, public health, and clinical research users. Shared tasks like the i2b2/VA Challenge address several of these barriers in part. Shared tasks provide annotated datasets to participants and sometimes to non-participants (i2b2 datasets are available to others a year after the Challenge). The i2b2 shared task is standardizing its corpus as much as possible—the same records are used from one year to the next with layers of annotation that build on each other, and common input/output specifications are applied every year. Shared tasks partially address the barrier of reproducibility by providing an evaluation opportunity that minimizes the risk of over-fitting: participants have time to train their systems in supervised fashion with an annotated training dataset, but then evaluation must be performed against a separate non-annotated dataset within a stringent time limit that prevents non-trivial system modifications. Although shared tasks are not designed for this purpose, the i2b2 Challenge has been the impetus for some new collaborations across independent research groups. Shared tasks have driven progress in related fields. For example, progress in speech understanding research was driven by a series of evaluations funded by DARPA from the late 1980s to the early 2000s.21 The research community was able to consistently drive down the error rate by a factor of two every 2 years, on successively more challenging tasks, moving from recognition of small-vocabulary read speech to automated transcription of broadcast news in multiple languages. Associated with this progress was the incorporation of speech recognition products into applications, from dictation to speech interfaces. Shared tasks provide value to the NLP community in several ways: Common evaluation metrics are developed. Annotated datasets are made available. Enticed by available annotated datasets, researchers in overlapping fields (both academic and corporate) participate in the tasks, bringing in new people and new approaches. Benchmarking evaluation on a shared dataset reveals the state-of-the-art performance for a given task. Students and post-docs receive excellent training opportunities. Preliminary results can be obtained by a new research group, which can potentially lead to funding opportunities. Pre-processed, standardized corpora with multiple layers of annotations on the same corpus pave the way for end-to-end evaluations in addition to evaluation on a single annotation layer. Conventions for standardizing annotations and input/output formats are developed, and despite other standardization efforts, shared task corpora often set de facto standards. In spite of the value of shared tasks, the tasks have several shortcomings: Participants come mainly from teams with funded projects that overlap with the shared task. For academic participants, a significant motivation is the opportunity to publish; however, there is sometimes limited value for the larger community in publications resulting from a shared task. Because development time is limited during shared tasks, participants often build on applications that already exist and apply methods already described in the literature. This can result in many similar approaches being applied to the same task. Although publishing the high-performing systems can be interesting, the resulting publications may not be novel and therefore may not improve the general body of knowledge. If a particular challenge task is repeated over time, there is a tendency for system approaches to converge on the approach that showed most success in the previous evaluation—evaluations repeated over time tend to reduce the diversity of approaches. Although shared tasks contribute to growth and progress, increased benefit to the community of clinical NLP developers and to potential users will require additional individual and community efforts that target existing barriers creatively. Driving progress in a way that will increase the impact of NLP in the realm of individual and population health will require creativity at both the grass-roots and the community levels. No single activity can tackle all barriers. In addition to encouraging variations on the development of shared tasks and their incentives, we would like to see new types of shared activities that foster the outcomes described below. In exchange for access to the costly annotated dataset, shared task participation could be contingent on depositing code in a shared repository or creating a web service for prospective users. In this model, the organizers could send test data to the participants' servers and the servers return the results for evaluation (see Leitner et al22 for a description of a metaserver used in evaluation of results from BioCreative II). The servers (and metaserver) could even persist, providing services to interested users beyond the initial shared task. Publication of computational methods in biomedical informatics journals like JAMIA could further encourage reproducibility of results through policies mandating simultaneous submission of code with a manuscript, as recommended by Pedersen.16 The clinical NLP community could independently accelerate reproducibility (and lead by example) by depositing code and developing web services in a common repository.23 This trend is occurring in settings restricted by affiliation, such as the VA VINCI framework24 (available only to VA researchers) and a cloud environment being hosted by the SHARP Area 4 grant18 (available to grant participants). The new National Center for Biomedical Computing iDASH25 is developing a similar cyber-infrastructure that will be restricted not by affiliation but by adherence to privacy policies and agreements required by data contributors. The National Library of Medicine is currently hosting a registry developed by the AMIA NLP working group called ORBIT for listing and pointing to biomedical informatics and NLP resources.26 In a shared task, dozens of research groups duplicate the same task independently. Although a variety of techniques can emerge for the same task, given the relatively short time frame allowed for development and training, the features and approaches applied in the challenge are often very similar. Whereas similarity of approaches reveals agreement among teams on the best approaches and sets the stage for collaboration, the competitive nature of shared tasks provides a disincentive to collaboration; the reward system for shared tasks is not at all dependent on the ability to collaborate across teams but is solely geared toward competition in which a single winner arises. The open source development community has found inherent rewards in collaborative development through supportive environments like GitHub. Perhaps we can learn from collaborative development communities who participate in hackathons27 and from the games industry in asking how a shared task can be designed so that collaboration is rewarded and becomes worthwhile, interesting, and attractive.28 Evaluation of a shared task is focused on accuracy, and existing challenges evaluate only predictive performance, not software engineering characteristics or usability. Imagine a shared task in which success is judged on usability of a system or direct portability of one technique to a new task or domain. Evaluating success of a system with this paradigm is inherently more complex, but we could learn from groupware evaluation, from the rich field of usability testing, and from incentivized competitions like those sponsored by the X-Prize Foundation.29 Because of the cost of creating annotated training data, shared tasks are often small scale, at least relative to real medical applications. We need new approaches to rapid adaptation of NLP systems to new applications, with less dependence on ‘deeply annotated’ data; such applications would present important opportunities for collaboration with the end user community, who might be motivated to provide domain expertise if they were likely to get a scalable, maintainable system out of the collaboration. Scalability will require more efficient techniques for manual annotation. And scalability will require an enriched ability to produce high quality software, which may necessitate better collaboration with industry30 and funding models that include support for operational development. The shared i2b2 evaluations have made a huge contribution to stimulating and vitalizing the field of clinical NLP; however, to ensure the transition into usable applications, the clinical NLP research community needs to address the critical issues of data access, development of shared infrastructure, and integration of software engineering methods to ensure the usability, maintainability, and availability of clinical NLP tools that are integrated into the workflow of real biomedical applications. This must be done in close collaboration with end users, software engineers, and clinical practitioners. We as a community need to think beyond the status quo of incremental improvement in the F score toward imaginative approaches that encourage collaboration, promote reproducibility, increase the scalability of NLP development, and provide value to end users. Authors are funded in part by U54HL108460, R01GM090187, SHARP ONC award 90TR0002, R01 CA127979, and U54LM008748. None. Not commissioned; internally peer reviewed. Wendy W. Chapman, Prakash M. Nadkarni, Lynette Hirschman, Leonard W. D'Avolio, Guergana K. Savova, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 5 |
| 2011 | Anaphoric relations in the clinical narrative: corpus creationabstractOBJECTIVE: The long-term goal of this work is the automated discovery of anaphoric relations from the clinical narrative. The creation of a gold standard set from a cross-institutional corpus of clinical notes and high-level characteristics of that gold standard are described. METHODS: A standard methodology for annotation guideline development, gold standard annotations, and inter-annotator agreement (IAA) was used. RESULTS: The gold standard annotations resulted in 7214 markables, 5992 pairs, and 1304 chains. Each report averaged 40 anaphoric markables, 33 pairs, and seven chains. The overall IAA is high on the Mayo dataset (0.6607), and moderate on the University of Pittsburgh Medical Center (UPMC) dataset (0.4072). The IAA between each annotator and the gold standard is high (Mayo: 0.7669, 0.7697, and 0.9021; UPMC: 0.6753 and 0.7138). These results imply a quality corpus feasible for system development. They also suggest the complementary nature of the annotations performed by the experts and the importance of an annotator team with diverse knowledge backgrounds. LIMITATIONS: Only one of the annotators had the linguistic background necessary for annotation of the linguistic attributes. The overall generalizability of the guidelines will be further strengthened by annotations of data from additional sites. This will increase the overall corpus size and the representation of each relation type. CONCLUSION: The first step toward the development of an anaphoric relation resolver as part of a comprehensive natural language processing system geared specifically for the clinical narrative in the electronic medical record is described. The deidentified annotated corpus will be available to researchers. Guergana K. Savova, Wendy W. Chapman, Jiaping Zheng, Rebecca S. Jacobson |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Drug side effect extraction from clinical narratives of psychiatry and psychology patientsabstractOBJECTIVE: To extract physician-asserted drug side effects from electronic medical record clinical narratives. MATERIALS AND METHODS: Pattern matching rules were manually developed through examining keywords and expression patterns of side effects to discover an individual side effect and causative drug relationship. A combination of machine learning (C4.5) using side effect keyword features and pattern matching rules was used to extract sentences that contain side effect and causative drug pairs, enabling the system to discover most side effect occurrences. Our system was implemented as a module within the clinical Text Analysis and Knowledge Extraction System. RESULTS: The system was tested in the domain of psychiatry and psychology. The rule-based system extracting side effects and causative drugs produced an F score of 0.80 (0.55 excluding allergy section). The hybrid system identifying side effect sentences had an F score of 0.75 (0.56 excluding allergy section) but covered more side effect and causative drug pairs than individual side effect extraction. DISCUSSION: The rule-based system was able to identify most side effects expressed by clear indication words. More sophisticated semantic processing is required to handle complex side effect descriptions in the narrative. We demonstrated that our system can be trained to identify sentences with complex side effect descriptions that can be submitted to a human expert for further abstraction. CONCLUSION: Our system was able to extract most physician-asserted drug side effects. It can be used in either an automated mode for side effect extraction or semi-automated mode to identify side effect sentences that can significantly simplify abstraction by a human expert. Sunghwan Sohn, Jean-Pierre A. Kocher, Christopher G. Chute, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 4 |
| 2011 | Coreference resolution: A review of general methodologies and applications in the clinical domain
Jiaping Zheng, Wendy W. Chapman, Rebecca S. Jacobson, Guergana K. Savova |
J. Biomed. Informatics | 4 |
| 2010 | Time-Oriented Question Answering from Clinical Narratives Using Semantic-Web Techniques
Cui Tao, Harold R. Solbrig, Deepak K. Sharma, Wei-Qi Wei, Guergana K. Savova, Christopher G. Chute |
ISWC (2) | 5 |
| 2010 | Leveraging informatics for genetic studies: use of the electronic medical record to enable a genome-wide association study of peripheral arterial diseaseabstractBACKGROUND: There is significant interest in leveraging the electronic medical record (EMR) to conduct genome-wide association studies (GWAS). METHODS: A biorepository of DNA and plasma was created by recruiting patients referred for non-invasive lower extremity arterial evaluation or stress ECG. Peripheral arterial disease (PAD) was defined as a resting/post-exercise ankle-brachial index (ABI) less than or equal to 0.9, a history of lower extremity revascularization, or having poorly compressible leg arteries. Controls were patients without evidence of PAD. Demographic data and laboratory values were extracted from the EMR. Medication use and smoking status were established by natural language processing of clinical notes. Other risk factors and comorbidities were ascertained based on ICD-9-CM codes, medication use and laboratory data. RESULTS: Of 1802 patients with an abnormal ABI, 115 had non-atherosclerotic vascular disease such as vasculitis, Buerger's disease, trauma and embolism (phenocopies) based on ICD-9-CM diagnosis codes and were excluded. The PAD cases (66+/-11 years, 64% men) were older than controls (61+/-8 years, 60% men) but had similar geographical distribution and ethnic composition. Among PAD cases, 1444 (85.6%) had an abnormal ABI, 233 (13.8%) had poorly compressible arteries and 10 (0.6%) had a history of lower extremity revascularization. In a random sample of 95 cases and 100 controls, risk factors and comorbidities ascertained from EMR-based algorithms had good concordance compared with manual record review; the precision ranged from 67% to 100% and recall from 84% to 100%. CONCLUSION: This study demonstrates use of the EMR to ascertain phenocopies, phenotype heterogeneity and relevant covariates to enable a GWAS of PAD. Biorepositories linked to EMR may provide a relatively efficient means of conducting GWAS. Iftikhar J. Kullo, Jyotishman Pathak, Guergana K. Savova, Zeenat Ali, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 4 |
| 2010 | Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applicationsabstractWe aim to build and evaluate an open-source natural language processing system for information extraction from electronic medical record clinical free-text. We describe and evaluate our system, the clinical Text Analysis and Knowledge Extraction System (cTAKES), released open-source at http://www.ohnlp.org. The cTAKES builds on existing open-source technologies-the Unstructured Information Management Architecture framework and OpenNLP natural language processing toolkit. Its components, specifically trained for the clinical domain, create rich linguistic and semantic annotations. Performance of individual components: sentence boundary detector accuracy=0.949; tokenizer accuracy=0.949; part-of-speech tagger accuracy=0.936; shallow parser F-score=0.924; named entity recognizer and system-level evaluation F-score=0.715 for exact and 0.824 for overlapping spans, and accuracy for concept mapping, negation, and status attributes for exact and overlapping spans of 0.957, 0.943, 0.859, and 0.580, 0.939, and 0.839, respectively. Overall performance is discussed against five applications. The cTAKES annotations are the foundation for methods and modules for higher-level semantic processing of clinical free-text. Guergana K. Savova, James J. Masanz, Philip V. Ogren, Jiaping Zheng, Sunghwan Sohn, Karin Kipper Schuler, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Towards Temporal Relation Discovery from the Clinical Narrative
Guergana K. Savova, Steven Bethard, William F. Styler IV, James H. Martin, Martha Palmer, James J. Masanz, Wayne H. Ward |
AMIA | 1 |
| 2009 | Mayo Clinic Smoking Status Classification System: Extensions and Improvements
Sunghwan Sohn, Guergana K. Savova |
AMIA | 2 |
| 2009 | Automatically extracting cancer disease characteristics from pathology reports into a Disease Knowledge Representation Model
Anni Coden, Guergana K. Savova, Igor L. Sominsky, Michael A. Tanenblatt, James J. Masanz, Karin Kipper Schuler, James W. Cooper, Piet C. de Groen |
J. Biomed. Informatics | 2 |
| 2008 | Constructing Evaluation Corpora for Automated Clinical Named Entity Recognition
Philip V. Ogren, Guergana K. Savova, Christopher G. Chute |
LREC | 2 |
| 2008 | System Evaluation on a Named Entity Corpus from Clinical Notes
Karin Kipper Schuler, Vinod Kaggal, James J. Masanz, Philip V. Ogren, Guergana K. Savova |
LREC | 5 |
| 2008 | Technical Brief: Mayo Clinic NLP System for Patient Smoking Status IdentificationabstractThis article describes our system entry for the 2006 I2B2 contest "Challenges in Natural Language Processing for Clinical Data" for the task of identifying the smoking status of patients. Our system makes the simplifying assumption that patient-level smoking status determination can be achieved by accurately classifying individual sentences from a patient's record. We created our system with reusable text analysis components built on the Unstructured Information Management Architecture and Weka. This reuse of code minimized the development effort related specifically to our smoking status classifier. We report precision, recall, F-score, and 95% exact confidence intervals for each metric. Recasting the classification task for the sentence level and reusing code from other text analysis projects allowed us to quickly build a classification system that performs with a system F-score of 92.64 based on held-out data tests and of 85.57 on the formal evaluation data. Our general medical natural language engine is easily adaptable to a real-world medical informatics application. Some of the limitations as applied to the use-case are negation detection and temporal resolution. Guergana K. Savova, Philip V. Ogren, Patrick H. Duffy, James D. Buntrock, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | Word sense disambiguation across two domains: Biomedical literature and clinical notes
Guergana K. Savova, Anni Coden, Igor L. Sominsky, Rie Johnson, Philip V. Ogren, Piet C. de Groen, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2006 | Building and Evaluating Annotated Corpora for Medical NLP Systems
Philip V. Ogren, Guergana K. Savova, James D. Buntrock, Christopher G. Chute |
AMIA | 2 |
| 2005 | Frame Semantics and the Domain of Functioning, Disability and Health
Guergana K. Savova, Marcelline R. Harris, Serguei V. S. Pakhomov, Christopher G. Chute |
AMIA | 1 |
| 2003 | Testing the Generalizability of the ISO Model for Nursing Diagnoses
Marcelline R. Harris, Hyeon-Eui Kim, Lori Rhudy, Guergana K. Savova, Christopher G. Chute |
AMIA | 4 |
| 2003 | A Data-Driven Approach for Extracting "the Most Specific Term" for Ontology Development
Guergana K. Savova, Marcelline R. Harris, Thomas M. Johnson, Serguei V. S. Pakhomov, Christopher G. Chute |
AMIA | 1 |
| 2003 | Designing for errors: similarities and differences of disfluency rates and prosodic characteristics across domains
Guergana K. Savova, Joan Bachenko |
INTERSPEECH | 1 |
| 2003 | A term extraction tool for expanding content in the domain of functioning, disability, and health: proof of concept
Marcelline R. Harris, Guergana K. Savova, Thomas M. Johnson, Christopher G. Chute |
J. Biomed. Informatics | 2 |
| 2000 | Improving language model perplexity and recognition accuracy for medical dictations via within-domain interpolation with literal and semi-literal corporaabstractWe propose a technique for improving language modeling for automated speech recognition of medical dictations by interpolating finished text (25M words) with small human-generated literal or/and machine-generated semiliteral corpora. By building and testing interpolated (ILM) with literal (LILM), semiliteral (SILM) and partial (PILM) corpora, we show that both perplexity and recognition results improve significantly with LILM and SILM; the two yielding very close results. 1. Guergana K. Savova, Michael Schonwetter, Serguei V. S. Pakhomov |
INTERSPEECH | 1 |