VLDB 2026 Research / reviewers in the wild / expert
Özlem Uzuner
dblp:43/957
· DBLP profile ↗
90ranked-venue papers
18as first author
28since 2021 · last 2026
0000-0001-8011-9850ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 76 · 13 first-author · 20 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM ApproachesabstractLarge language models (LLMs) have shown considerable promise in clinical natural language processing, yet few domain-specific datasets exist to rigorously evaluate their performance on radiology tasks. In this work, we introduce an annotated corpus of 6,393 radiology reports from 586 patients, each labeled for follow-up imaging status, to support the development and benchmarking of follow-up adherence detection systems. Using this corpus, we systematically compared traditional machine-learning classifiers, including logistic regression (LR), support vector machines (SVM), Longformer, and a fully fine-tuned Llama3-8B-Instruct, with recent generative LLMs. To evaluate generative LLMs, we tested GPT-4o and the open-source GPT-OSS-20B under two configurations: a baseline (Base) and a task-optimized (Advanced) setting that focused inputs on metadata, recommendation sentences, and their surrounding context. A refined prompt for GPT-OSS-20B further improved reasoning accuracy. Performance was assessed using precision, recall, and F1 scores with 95% confidence intervals estimated via non-parametric bootstrapping. Inter-annotator agreement was high (F1 = 0.846). GPT-4o (Advanced) achieved the best performance (F1 = 0.832), followed closely by GPT-OSS-20B (Advanced; F1 = 0.828). LR and SVM also performed strongly (F1 = 0.776 and 0.775), underscoring that while LLMs approach human-level agreement through prompt optimization, interpretable and resource-efficient models remain valuable baselines. Namu Park, Giridhar Kaushik Ramachandran, Kevin Lybarger, Fei Xia 0004, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
LREC | 5 |
| 2026 | Mental Health Disorder Detection beyond Social Media: A Systematic Review of Available Datasets
Sadiya Sayara Chowdhury Puspo, Ana-Maria Bucur, Stevie Chancellor, Özlem Uzuner, Marcos Zampieri |
LREC | 4 |
| 2026 | Automated identification of incidentalomas requiring follow-up: A multi-anatomy evaluation of LLM-based and supervised approaches
Namu Park, Farzad Ahmed, Zhaoyi Sun, Kevin Lybarger, Ethan Breinhorst, Julie Hu, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
J. Biomed. Informatics | 7 |
| 2025 | A scoping review of natural language processing in addressing medically inaccurate information: Errors, misinformation, and hallucination
Zhaoyi Sun, Wen-Wai Yim, Özlem Uzuner, Fei Xia 0004, Meliha Yetisgen |
J. Biomed. Informatics | 3 |
| 2024 | Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and MethodsabstractSocial determinants of health (SDoH) play a critical role in shaping health outcomes, particularly in pediatric populations where interventions can have long-term implications. SDoH are frequently studied in the Electronic Health Record (EHR), which provides a rich repository for diverse patient data. In this work, we present a novel annotated corpus, the Pediatric Social History Annotation Corpus (PedSHAC), and evaluate the automatic extraction of detailed SDoH representations using fine-tuned and in-context learning methods with Large Language Models (LLMs). PedSHAC comprises annotated social history sections from 1,260 clinical notes obtained from pediatric patients within the University of Washington (UW) hospital system. Employing an event-based annotation scheme, PedSHAC captures ten distinct health determinants to encompass living and economic stability, prior trauma, education access, substance use history, and mental health with an overall annotator agreement of 81.9 F1. Our proposed fine-tuning LLM-based extractors achieve high performance at 78.4 F1 for event arguments. In-context learning approaches with GPT-4 demonstrate promise for reliable SDoH extraction with limited annotated examples, with extraction performance at 82.3 F1 for event triggers. Yujuan Fu, Giridhar Kaushik Ramachandran, Nicholas J. Dobbins, Namu Park, Michael Leu, Abby R. Rosenberg, Kevin Lybarger, Fei Xia 0004, Özlem Uzuner, Meliha Yetisgen |
LREC/COLING | 9 |
| 2024 | A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language ModelsabstractMedical imaging is critical to the diagnosis, surveillance, and treatment of many health conditions, including oncological, neurological, cardiovascular, and musculoskeletal disorders, among others. Radiologists interpret these complex, unstructured images and articulate their assessments through narrative reports that remain largely unstructured. This unstructured narrative must be converted into a structured semantic representation to facilitate secondary applications such as retrospective analyses or clinical decision support. Here, we introduce the Corpus of Annotated Medical Imaging Reports (CAMIR), which includes 609 annotated radiology reports from three imaging modality types: Computed Tomography, Magnetic Resonance Imaging, and Positron Emission Tomography-Computed Tomography. Reports were annotated using an event-based schema that captures clinical indications, lesions, and medical problems. Each event consists of a trigger and multiple arguments, and a majority of the argument types, including anatomy, normalize the spans to pre-defined concepts to facilitate secondary use. CAMIR uniquely combines a granular event structure and concept normalization. To extract CAMIR events, we explored two BERT (Bi-directional Encoder Representation from Transformers)-based architectures, including an existing architecture (mSpERT) that jointly extracts all event information and a multi-step approach (PL-Marker++) that we augmented for the CAMIR schema. Namu Park, Kevin Lybarger, Giridhar Kaushik Ramachandran, Spencer Lewis, Aashka Damani, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
LREC/COLING | 6 |
| 2024 | CACER: Clinical concept Annotations for Cancer Events and RelationsabstractOBJECTIVE: Clinical notes contain unstructured representations of patient histories, including the relationships between medical problems and prescription drugs. To investigate the relationship between cancer drugs and their associated symptom burden, we extract structured, semantic representations of medical problem and drug information from the clinical narratives of oncology notes. MATERIALS AND METHODS: We present Clinical concept Annotations for Cancer Events and Relations (CACER), a novel corpus with fine-grained annotations for over 48 000 medical problems and drug events and 10 000 drug-problem and problem-problem relations. Leveraging CACER, we develop and evaluate transformer-based information extraction models such as Bidirectional Encoder Representations from Transformers (BERT), Fine-tuned Language Net Text-To-Text Transfer Transformer (Flan-T5), Large Language Model Meta AI (Llama3), and Generative Pre-trained Transformers-4 (GPT-4) using fine-tuning and in-context learning (ICL). RESULTS: In event extraction, the fine-tuned BERT and Llama3 models achieved the highest performance at 88.2-88.0 F1, which is comparable to the inter-annotator agreement (IAA) of 88.4 F1. In relation extraction, the fine-tuned BERT, Flan-T5, and Llama3 achieved the highest performance at 61.8-65.3 F1. GPT-4 with ICL achieved the worst performance across both tasks. DISCUSSION: The fine-tuned models significantly outperformed GPT-4 in ICL, highlighting the importance of annotated training data and model optimization. Furthermore, the BERT models performed similarly to Llama3. For our task, large language models offer no performance advantage over the smaller BERT models. CONCLUSIONS: We introduce CACER, a novel corpus with fine-grained annotations for medical problems, drugs, and their relationships in clinical narratives of oncology notes. State-of-the-art transformer models achieved performance comparable to IAA for several extraction tasks. Yujuan Fu, Giridhar Kaushik Ramachandran, Ahmad Halwani, Bridget T. McInnes, Fei Xia 0004, Kevin Lybarger, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | Large language models for biomedicine: foundations, opportunities, challenges, and best practicesabstractOBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications. Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Clinical natural language processing for secondary uses
Yanjun Gao, Diwakar Mahajan, Özlem Uzuner, Meliha Yetisgen |
J. Biomed. Informatics | 3 |
| 2023 | LeafAI: query generator for clinical cohort discovery rivaling a human programmerabstractOBJECTIVE: Identifying study-eligible patients within clinical databases is a critical step in clinical research. However, accurate query design typically requires extensive technical and biomedical expertise. We sought to create a system capable of generating data model-agnostic queries while also providing novel logical reasoning capabilities for complex clinical trial eligibility criteria. MATERIALS AND METHODS: The task of query creation from eligibility criteria requires solving several text-processing problems, including named entity recognition and relation extraction, sequence-to-sequence transformation, normalization, and reasoning. We incorporated hybrid deep learning and rule-based modules for these, as well as a knowledge base of the Unified Medical Language System (UMLS) and linked ontologies. To enable data-model agnostic query creation, we introduce a novel method for tagging database schema elements using UMLS concepts. To evaluate our system, called LeafAI, we compared the capability of LeafAI to a human database programmer to identify patients who had been enrolled in 8 clinical trials conducted at our institution. We measured performance by the number of actual enrolled patients matched by generated queries. RESULTS: LeafAI matched a mean 43% of enrolled patients with 27 225 eligible across 8 clinical trials, compared to 27% matched and 14 587 eligible in queries by a human database programmer. The human programmer spent 26 total hours crafting queries compared to several minutes by LeafAI. CONCLUSIONS: Our work contributes a state-of-the-art data model-agnostic query generation system capable of conditional reasoning using a knowledge base. We demonstrate that LeafAI can rival an experienced human programmer in finding patients eligible for clinical trials. Nicholas J. Dobbins, Weipeng Zhou, Kristine Lan, H. Nina Kim, Robert D. Harrington, Özlem Uzuner, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Leveraging natural language processing to augment structured social determinants of health data in the electronic health recordabstractOBJECTIVE: Social determinants of health (SDOH) impact health outcomes and are documented in the electronic health record (EHR) through structured data and unstructured clinical notes. However, clinical notes often contain more comprehensive SDOH information, detailing aspects such as status, severity, and temporality. This work has two primary objectives: (1) develop a natural language processing information extraction model to capture detailed SDOH information and (2) evaluate the information gain achieved by applying the SDOH extractor to clinical narratives and combining the extracted representations with existing structured data. MATERIALS AND METHODS: We developed a novel SDOH extractor using a deep learning entity and relation extraction architecture to characterize SDOH across various dimensions. In an EHR case study, we applied the SDOH extractor to a large clinical data set with 225 089 patients and 430 406 notes with social history sections and compared the extracted SDOH information with existing structured data. RESULTS: The SDOH extractor achieved 0.86 F1 on a withheld test set. In the EHR case study, we found extracted SDOH information complements existing structured data with 32% of homeless patients, 19% of current tobacco users, and 10% of drug users only having these health risk factors documented in the clinical narrative. CONCLUSIONS: Utilizing EHR data to identify SDOH health risk factors and social needs may improve patient care and outcomes. Semantic representations of text-encoded SDOH information can augment existing structured data, and this more comprehensive SDOH representation can assist health systems in identifying and addressing these social needs. Kevin Lybarger, Nicholas J. Dobbins, Ritche Long, Angad P. Singh, Patrick Wedgeworth, Özlem Uzuner, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Advancements in extracting social determinants of health information from narrative textabstractSocial determinants of health (SDoH) are the conditions in which people are born, live, work, and age that affect personal well-being, health outcomes, and life expectancy.1 SDoH include a range of nonmedical factors, including substance use, quality of domestic life, marital status, employment status, education, race, geography, and other factors that impact health. Understanding patient SDoH can inform patient health care and has the potential to improve health outcomes and reduce health disparities.2,3 Patient SDoH information is documented in the electronic health record (EHR) and other health-related databases through structured data and free-text (natural language) documents, including patient notes. For many SDoH, the free-text descriptions capture social and behavioral factors with higher prevalence and more detail than is available through structured data. Utilizing free-text SDoH information in large-scale studies, clinical decision-support systems, and other secondary use applications, requires the automatic extraction of key aspects of the SDoH using natural language processing (NLP). NLP-based information extraction maps the unstructured, free-text descriptions of SDoH to structured semantic representations that can be combined with available structured data to create more complete patient profiles. Kevin Lybarger, Oliver J. Bear Don't Walk IV, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | The 2022 n2c2/UW shared task on extracting social determinants of healthabstractOBJECTIVE: The n2c2/UW SDOH Challenge explores the extraction of social determinant of health (SDOH) information from clinical notes. The objectives include the advancement of natural language processing (NLP) information extraction techniques for SDOH and clinical information more broadly. This article presents the shared task, data, participating teams, performance results, and considerations for future work. MATERIALS AND METHODS: The task used the Social History Annotated Corpus (SHAC), which consists of clinical text with detailed event-based annotations for SDOH events, such as alcohol, drug, tobacco, employment, and living situation. Each SDOH event is characterized through attributes related to status, extent, and temporality. The task includes 3 subtasks related to information extraction (Subtask A), generalizability (Subtask B), and learning transfer (Subtask C). In addressing this task, participants utilized a range of techniques, including rules, knowledge bases, n-grams, word embeddings, and pretrained language models (LM). RESULTS: A total of 15 teams participated, and the top teams utilized pretrained deep learning LM. The top team across all subtasks used a sequence-to-sequence approach achieving 0.901 F1 for Subtask A, 0.774 F1 Subtask B, and 0.889 F1 for Subtask C. CONCLUSIONS: Similar to many NLP tasks and domains, pretrained LM yielded the best performance, including generalizability and learning transfer. An error analysis indicates extraction performance varies by SDOH, with lower performance achieved for conditions, like substance use and homelessness, which increase health risks (risk factors) and higher performance achieved for conditions, like substance abstinence and living with family, which reduce health risks (protective factors). Kevin Lybarger, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Progress Note Understanding - Assessment and Plan Reasoning: Overview of the 2022 N2C2 Track 3 shared task
Yanjun Gao, Dmitriy Dligach, Timothy A. Miller, Matthew M. Churpek, Özlem Uzuner, Majid Afshar |
J. Biomed. Informatics | 5 |
| 2023 | Overview of the 2022 n2c2 shared task on contextualized medication event extraction in clinical notesabstractBACKGROUND: An accurate medication history, foundational for providing quality medical care, requires understanding of medication change events documented in clinical notes. However, extracting medication changes without the necessary clinical context is insufficient for real-world applications. METHODS: To address this need, Track 1 of the 2022 National NLP Clinical Challenges focused on extracting the context for medication changes documented in clinical notes using the Contextualized Medication Event Dataset. Track 1 consisted of 3 subtasks: extracting medication mentions from clinical notes (NER), determining whether a medication change is being discussed (Event), and determining the action, negation, temporality, certainty, and actor for any change events (Context). Participants were allowed to participate in any one or more of the subtasks. RESULTS: A total of 32 teams with participants from 19 countries submitted a total of 211 systems across all subtasks. Most teams formulated NER as a token classification task and Event and Context as multi-class classification tasks, using transformer-based large language models. Overall, performance for NER was high across submitted systems. However, performance for Event and Context were much lower, often due to indirectly stated change events with no clear action verb, events requiring farther textual clues for understanding, and medication mentions with multiple change events. CONCLUSIONS: This shared task showed that while NLP research on medication extraction is relatively mature, understanding of contextual information surrounding medication events in clinical notes is still an open problem requiring further research to achieve the end goal of supporting real-world clinical applications. Diwakar Mahajan, Jennifer J. Liang, Ching-Huei Tsou, Özlem Uzuner |
J. Biomed. Informatics | 4 |
| 2023 | Extracting medication changes in clinical narratives using pre-trained language models
Giridhar Kaushik Ramachandran, Kevin Lybarger, Yaya Liu, Diwakar Mahajan, Jennifer J. Liang, Ching-Huei Tsou, Meliha Yetisgen, Özlem Uzuner |
J. Biomed. Informatics | 8 |
| 2022 | A Corpus of Radiology Reports From Multiple Imaging Modalities With Fine-grained Event-based Annotations
Kevin Lybarger, Namu Park, Sitong Zhou, Aashka Damani, Alison Brennan, Jagjeet Gill, Nianiella Dorvall, Vy Huynh, Spencer Lewis, Martin L. Gunn, Özlem Uzuner, Meliha Yetisgen |
AMIA | 11 |
| 2022 | A Context-Enhanced De-identification SystemabstractMany modern entity recognition systems, including the current state-of-the-art de-identification systems, are based on bidirectional long short-term memory (biLSTM) units augmented by a conditional random field (CRF) sequence optimizer. These systems process the input sentence by sentence. This approach prevents the systems from capturing dependencies over sentence boundaries and makes accurate sentence boundary detection a prerequisite. Since sentence boundary detection can be problematic especially in clinical reports, where dependencies and co-references across sentence boundaries are abundant, these systems have clear limitations. In this study, we built a new system on the framework of one of the current state-of-the-art de-identification systems, NeuroNER, to overcome these limitations. This new system incorporates context embeddings through forward and backward n -grams without using sentence boundaries. Our context-enhanced de-identification (CEDI) system captures dependencies over sentence boundaries and bypasses the sentence boundary detection problem altogether. We enhanced this system with deep affix features and an attention mechanism to capture the pertinent parts of the input. The CEDI system outperforms NeuroNER on the 2006 i2b2 de-identification challenge dataset, the 2014 i2b2 shared task de-identification dataset, and the 2016 CEGS N-GRID de-identification dataset ( p < 0.01 ). All datasets comprise narrative clinical reports in English but contain different note types varying from discharge summaries to psychiatric notes. Enhancing CEDI with deep affix features and the attention mechanism further increased performance. Kahyun Lee, Mehmet Kayaalp 0002, Sam Henry 0001, Özlem Uzuner |
ACM Trans. Comput. Heal. | 4 |
| 2022 | A scoping review of publicly available language tasks in clinical natural language processingabstractOBJECTIVE: To provide a scoping review of papers on clinical natural language processing (NLP) shared tasks that use publicly available electronic health record data from a cohort of patients. MATERIALS AND METHODS: We searched 6 databases, including biomedical research and computer science literature databases. A round of title/abstract screening and full-text screening were conducted by 2 reviewers. Our method followed the PRISMA-ScR guidelines. RESULTS: A total of 35 papers with 48 clinical NLP tasks met inclusion criteria between 2007 and 2021. We categorized the tasks by the type of NLP problems, including named entity recognition, summarization, and other NLP tasks. Some tasks were introduced as potential clinical decision support applications, such as substance abuse detection, and phenotyping. We summarized the tasks by publication venue and dataset type. DISCUSSION: The breadth of clinical NLP tasks continues to grow as the field of NLP evolves with advancements in language systems. However, gaps exist with divergent interests between the general domain NLP community and the clinical informatics community for task motivation and design, and in generalizability of the data sources. We also identified issues in data preparation. CONCLUSION: The existing clinical NLP tasks cover a wide range of topics and the field is expected to grow and attract more attention from both general domain NLP and clinical informatics community. We encourage future work to incorporate multidisciplinary collaboration, reporting transparency, and standardization in data preparation. We provide a listing of all the shared task papers and datasets from this review in a GitLab repository. Yanjun Gao, Dmitriy Dligach, Leslie Christensen, Samuel Tesch, Ryan Laffin, Dongfang Xu, Timothy A. Miller, Özlem Uzuner, Matthew M. Churpek, Majid Afshar |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | Call for papers: Special issue on clinical natural language processing for secondary use applications
Meliha Yetisgen, Özlem Uzuner, Yanjun Gao, Diwakar Mahajan |
J. Biomed. Informatics | 2 |
| 2022 | Investigating the Effect of Preprocessing Arabic Text on Offensive Language and Hate Speech DetectionabstractPreprocessing of input text can play a key role in text classification by reducing dimensionality and removing unnecessary content. This study aims to investigate the impact of preprocessing on Arabic offensive language classification. We explore six preprocessing techniques: conversion of emojis to Arabic textual labels, normalization of different forms of Arabic letters, normalization of selected nouns from dialectal Arabic to Modern Standard Arabic, conversion of selected hyponyms to hypernyms, hashtag segmentation, and basic cleaning such as removing numbers, kashidas, diacritics, and HTML tags. We also experiment with raw text and a combination of all six preprocessing techniques. We apply different types of classifiers in our experiments including traditional machine learning, ensemble machine learning, Artificial Neural Networks, and Bidirectional Encoder Representations from Transformers (BERT)-based models to analyze the impact of preprocessing. Our results demonstrate significant variations in the effects of preprocessing on each classifier type and on each dataset. Classifiers that are based on BERT do not benefit from preprocessing, while traditional machine learning classifiers do. However, these results can benefit from validation on larger datasets that cover broader domains and dialects. Fatemah Husain, Özlem Uzuner |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2021 | A New Corpus for Clinical Findings in Radiology Reports
Wilson Lau, David Wayne, Spencer Lewis, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
AMIA | 4 |
| 2021 | Detecting Scarce Emotions Using BERT and Hyperparameter Optimization
Zahra Rajabi, Özlem Uzuner, Amarda Shehu |
ICANN (5) | 2 |
| 2021 | Scalable Graph Synthesis with Adj and 1 - AdjabstractGraph synthesis is a long-standing research problem.Many deep neural networks that learn about latent characteristics of graphs and generate fake graphs have been proposed.However, in many cases their scalability is too high to be used to synthesize large graphs.Recently, one work proposed an interesting scalable idea to learn and generate random walks that can be merged into a graph.Due to its difficulty, however, the random walk-based graph synthesis failed to show state-of-the-art performance in many cases.We present an improved random walk-based method by using negative random walks.In our experiments with 6 datasets and 8 baseline methods, our method shows the best performance in almost all cases.We achieve both high scalability and generation quality. Jinsung Jeon, Jing Liu 0024, Jayoung Kim 0002, Jaehoon Lee 0002, Noseong Park, Jamie Jooyeon Lee, Özlem Uzuner, Sushil Jajodia |
SDM | 7 |
| 2021 | Corrigendum to: The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task on clinical concept normalization for clinical recordsabstractObjective The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task track 3, focused on medical concept normalization (MCN) in clinical records. This track aimed to assess the state of the art in identifying and matching salient medical concepts to a controlled vocabulary. In this paper, we describe the task, describe the data set used, compare the participating systems, present results, identify the strengths and limitations of the current state of the art, and identify directions for future research. Materials and methods Participating teams were provided with narrative discharge summaries in which text spans corresponding to medical concepts were identified. This paper refers to these text spans as mentions. Teams were tasked with normalizing these mentions to concepts, represented by concept unique identifiers, within the Unified Medical Language System. Submitted systems represented 4 broad categories of approaches: cascading dictionary matching, cosine distance, deep learning, and retrieve-and-rank systems. Disambiguation modules were common across all approaches. Results A total of 33 teams participated in the MCN task. The best-performing team achieved an accuracy of 0.8526. The median and mean performances among all teams were 0.7733 and 0.7426, respectively. Conclusions Overall performance among the top 10 teams was high. However, several mention types were challenging for all teams. These included mentions requiring disambiguation of misspelled words, acronyms, abbreviations, and mentions with more than 1 possible semantic type. Also challenging were complex mentions of long, multi-word terms that may require new ways of extracting and representing mention meaning, the use of domain knowledge, parse trees, or hand-crafted rules. Sam Henry 0001, Yanshan Wang, Feichen Shen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Transferability of neural network clinical deidentification systemsabstractOBJECTIVE: Neural network deidentification studies have focused on individual datasets. These studies assume the availability of a sufficient amount of human-annotated data to train models that can generalize to corresponding test data. In real-world situations, however, researchers often have limited or no in-house training data. Existing systems and external data can help jump-start deidentification on in-house data; however, the most efficient way of utilizing existing systems and external data is unclear. This article investigates the transferability of a state-of-the-art neural clinical deidentification system, NeuroNER, across a variety of datasets, when it is modified architecturally for domain generalization and when it is trained strategically for domain transfer. MATERIALS AND METHODS: We conducted a comparative study of the transferability of NeuroNER using 4 clinical note corpora with multiple note types from 2 institutions. We modified NeuroNER architecturally to integrate 2 types of domain generalization approaches. We evaluated each architecture using 3 training strategies. We measured transferability from external sources; transferability across note types; the contribution of external source data when in-domain training data are available; and transferability across institutions. RESULTS AND CONCLUSIONS: Transferability from a single external source gave inconsistent results. Using additional external sources consistently yielded an F1-score of approximately 80%. Fine-tuning emerged as a dominant transfer strategy, with or without domain generalization. We also found that external sources were useful even in cases where in-domain training data were available. Transferability across institutions differed by note type and annotation label but resulted in improved performance. Kahyun Lee, Nicholas J. Dobbins, Bridget T. McInnes, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | MT-clinical BERT: scaling clinical information extraction with multitask learningabstractOBJECTIVE: Clinical notes contain an abundance of important, but not-readily accessible, information about patients. Systems that automatically extract this information rely on large amounts of training data of which there exists limited resources to create. Furthermore, they are developed disjointly, meaning that no information can be shared among task-specific systems. This bottleneck unnecessarily complicates practical application, reduces the performance capabilities of each individual solution, and associates the engineering debt of managing multiple information extraction systems. MATERIALS AND METHODS: We address these challenges by developing Multitask-Clinical BERT: a single deep learning model that simultaneously performs 8 clinical tasks spanning entity extraction, personal health information identification, language entailment, and similarity by sharing representations among tasks. RESULTS: We compare the performance of our multitasking information extraction system to state-of-the-art BERT sequential fine-tuning baselines. We observe a slight but consistent performance degradation in MT-Clinical BERT relative to sequential fine-tuning. DISCUSSION: These results intuitively suggest that learning a general clinical text representation capable of supporting multiple tasks has the downside of losing the ability to exploit dataset or clinical note-specific properties when compared to a single, task-specific model. CONCLUSIONS: We find our single system performs competitively with all state-the-art task-specific systems while also benefiting from massive computational benefits at inference. Andriy Mulyar, Özlem Uzuner, Bridget T. McInnes |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | A Survey of Offensive Language Detection for the Arabic LanguageabstractThe use of offensive language in user-generated content is a serious problem that needs to be addressed with the latest technology. The field of Natural Language Processing (NLP) can support the automatic detection of offensive language. In this survey, we review previous NLP studies that cover Arabic offensive language detection. This survey investigates the state-of-the-art in offensive language detection for the Arabic language, providing a structured overview of previous approaches, including core techniques, tools, resources, methods, and main features used. This work also discusses the limitations and gaps of the previous studies. Findings from this survey emphasize the importance of investing further effort in detecting Arabic offensive language, including the development of benchmark resources and the invention of novel preprocessing and feature extraction techniques. Fatemah Husain, Özlem Uzuner |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | Jointly Learning Clinical Entities and Relations with Contextual Language Models and Explicit Context
Paul Barry, Sam Henry 0001, Meliha Yetisgen, Bridget T. McInnes, Özlem Uzuner |
AMIA | 5 |
| 2020 | Medical Concept Normalization via Description Prediction
Kahyun Lee, Sam Henry 0001, Özlem Uzuner |
AMIA | 3 |
| 2020 | A Novel Corpus With Detailed Annotations of Social Determinants of Health
Kevin Lybarger, Kylie Kerker, Jolie Shen, Erica Qiao, Özlem Uzuner, Mari Ostendorf, Meliha Yetisgen |
AMIA | 6 |
| 2020 | 2018 n2c2 shared task on adverse drug events and medication extraction in electronic health recordsabstractOBJECTIVE: This article summarizes the preparation, organization, evaluation, and results of Track 2 of the 2018 National NLP Clinical Challenges shared task. Track 2 focused on extraction of adverse drug events (ADEs) from clinical records and evaluated 3 tasks: concept extraction, relation classification, and end-to-end systems. We perform an analysis of the results to identify the state of the art in these tasks, learn from it, and build on it. MATERIALS AND METHODS: For all tasks, teams were given raw text of narrative discharge summaries, and in all the tasks, participants proposed deep learning-based methods with hand-designed features. In the concept extraction task, participants used sequence labelling models (bidirectional long short-term memory being the most popular), whereas in the relation classification task, they also experimented with instance-based classifiers (namely support vector machines and rules). Ensemble methods were also popular. RESULTS: A total of 28 teams participated in task 1, with 21 teams in tasks 2 and 3. The best performing systems set a high performance bar with F1 scores of 0.9418 for concept extraction, 0.9630 for relation classification, and 0.8905 for end-to-end. However, the results were much lower for concepts and relations of Reasons and ADEs. These were often missed because local context is insufficient to identify them. CONCLUSIONS: This challenge shows that clinical concept extraction and relation classification systems have a high performance for many concept types, but significant improvement is still required for ADEs and Reasons. Incorporating the larger context or outside knowledge will likely improve the performance of future systems. Sam Henry 0001, Kevin Buchan, Michele Filannino, Amber Stubbs, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task on clinical concept normalization for clinical recordsabstractOBJECTIVE: The 2019 National Natural language processing (NLP) Clinical Challenges (n2c2)/Open Health NLP (OHNLP) shared task track 3, focused on medical concept normalization (MCN) in clinical records. This track aimed to assess the state of the art in identifying and matching salient medical concepts to a controlled vocabulary. In this paper, we describe the task, describe the data set used, compare the participating systems, present results, identify the strengths and limitations of the current state of the art, and identify directions for future research. MATERIALS AND METHODS: Participating teams were provided with narrative discharge summaries in which text spans corresponding to medical concepts were identified. This paper refers to these text spans as mentions. Teams were tasked with normalizing these mentions to concepts, represented by concept unique identifiers, within the Unified Medical Language System. Submitted systems represented 4 broad categories of approaches: cascading dictionary matching, cosine distance, deep learning, and retrieve-and-rank systems. Disambiguation modules were common across all approaches. RESULTS: A total of 33 teams participated in the MCN task. The best-performing team achieved an accuracy of 0.8526. The median and mean performances among all teams were 0.7733 and 0.7426, respectively. CONCLUSIONS: Overall performance among the top 10 teams was high. However, several mention types were challenging for all teams. These included mentions requiring disambiguation of misspelled words, acronyms, abbreviations, and mentions with more than 1 possible semantic type. Also challenging were complex mentions of long, multi-word terms that may require new ways of extracting and representing mention meaning, the use of domain knowledge, parse trees, or hand-crafted rules. Sam Henry 0001, Yanshan Wang, Feichen Shen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | Advancing the state of the art in automatic extraction of adverse drug events from narrativesabstractAdverse drug events (ADEs), defined as “any injuries resulting from medication use, including physical harm, mental harm, or loss of function,”1 are reported to account for approximately 30% of all adverse events,2 with results that can include repeated hospital admission and fatality. Information about causes of ADEs can be found in data that document concurrent use of multiple medications, drug interactions, and possible allergies such as “charts, laboratory [data], prescription data” and “administrative data.”3 However, much of the crucial information related to ADEs are detailed in free text narratives and are not easily accessible by computerized systems, requiring manual review and manual identification of this information. Natural language processing (NLP) holds potential for automatically extracting ADE-related information from narratives, to make it available for decision support systems that can alert clinicians to potential ADEs at the point of care. To assess and advance the state of the art in NLP for extraction of ADEs, the National NLP Clinical Challenges (n2c2) shared task in 2018 included a track on this topic.4 This track required the identification of potential ADE mentions, along with their link to the medication that caused them, and the administration details such as the dosage, route, and frequency information related to the medication causing the ADE. The systems that tackled extraction of ADEs and related concepts primarily utilized recurrent deep neural networks consisting of bidirectional long short-term memory units, achieving performances that reached 94% in F-measure. In linking ADEs to their causes, the systems were more diverse in their methods, utilizing a range of machine learning approaches including both deep learning and more traditional methods and achieving performances that reached 96% in F-measure. These results indicate that while they are not perfect, NLP systems can successfully extract ADE information from narratives with impressive accuracy. In this editorial, we highlight 4 systems. Others are summarized in Henry et al.4 Özlem Uzuner, Amber Stubbs, Leslie Lenert |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Adverse drug event detection using reason assignments in FDA drug labels
Corey Sutphin, Kahyun Lee, Antonio Jimeno-Yepes, Özlem Uzuner, Bridget T. McInnes |
J. Biomed. Informatics | 4 |
| 2019 | Automatic Identification of Social Determinants of Health from Clinical Records
Kevin Lybarger, Mari Ostendorf, Özlem Uzuner, Meliha Yetisgen |
AMIA | 3 |
| 2019 | Clinical Text Mining in Mental Health
Jessica D. Tenenbaum, Ramakanth Kavuluru, Thomas H. McCoy, Özlem Uzuner, Sumithra Velupillai |
AMIA | 4 |
| 2019 | Cohort selection for clinical trials: n2c2 2018 shared task track 1abstractOBJECTIVE: Track 1 of the 2018 National NLP Clinical Challenges shared tasks focused on identifying which patients in a corpus of longitudinal medical records meet and do not meet identified selection criteria. MATERIALS AND METHODS: To address this challenge, we annotated American English clinical narratives for 288 patients according to whether they met these criteria. We chose criteria from existing clinical trials that represented a variety of natural language processing tasks, including concept extraction, temporal reasoning, and inference. RESULTS: A total of 47 teams participated in this shared task, with 224 participants in total. The participants represented 18 countries, and the teams submitted 109 total system outputs. The best-performing system achieved a micro F1 score of 0.91 using a rule-based approach. The top 10 teams used rule-based and hybrid systems to approach the problems. DISCUSSION: Clinical narratives are open to interpretation, particularly in cases where the selection criterion may be underspecified. This leaves room for annotators to use domain knowledge and intuition in selecting patients, which may lead to error in system outputs. However, teams who consulted medical professionals while building their systems were more likely to have high recall for patients, which is preferable for patient selection systems. CONCLUSIONS: There is not yet a 1-size-fits-all solution for natural language processing systems approaching this task. Future research in this area can look to examining criteria requiring even more complex inferences, temporal reasoning, and domain knowledge. Amber Stubbs, Michele Filannino, Ergin Soysal, Sam Henry 0001, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | New approaches to cohort selectionabstractCohort selection for clinical trials is a critical component of modern medicine, yet it remains one of the most difficult, time-consuming, and expensive aspects of testing new treatments and interventions. Each clinical trial defines inclusion and exclusion criteria that describe the required patient population for the trial to accurately determine efficacy of the treatment. These criteria can be broad, limited only to specific ages or genders, or can be very specific, requiring certain medications be taken in a time period, or certain intentions on the parts of the patients (ie, an intention to become pregnant). While a simple database search can often identify patients of the right age, or even those with particular diagnoses or test results, the more complex criteria often require study staff to manually examine records to identify qualified patients. Or, studies may rely on the patients to seek out the trial, or to be directed to trial possibilities by their doctors—both of which may lead to representation bias and misleading conclusions for the trial.1,2 Amber Stubbs, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | Precision Medicine: matching cancer patients with clinical trials
Michele Filannino, Kahyun Lee, Di Jin 0005, Kevin Buchan, Willie Boag, Özlem Uzuner |
AMIA | 6 |
| 2018 | Precision Medicine: matching cancer patients with relevant treatments
Kahyun Lee, Michele Filannino, Di Jin 0005, Kevin Buchan, Willie Boag, Özlem Uzuner |
AMIA | 6 |
| 2018 | FABLE: A Semi-Supervised Prescription Information Extraction System
Carson Tao, Michele Filannino, Özlem Uzuner |
AMIA | 3 |
| 2018 | Extracting ADRs from Drug Labels using Bi-LSTM and CRFs
Carson Tao, Michele Filannino, Özlem Uzuner |
AMIA | 3 |
| 2018 | Segment convolutional neural networks (Seg-CNNs) for classifying relations in clinical notesabstractWe propose Segment Convolutional Neural Networks (Seg-CNNs) for classifying relations from clinical notes. Seg-CNNs use only word-embedding features without manual feature engineering. Unlike typical CNN models, relations between 2 concepts are identified by simultaneously learning separate representations for text segments in a sentence: preceding, concept1, middle, concept2, and succeeding. We evaluate Seg-CNN on the i2b2/VA relation classification challenge dataset. We show that Seg-CNN achieves a state-of-the-art micro-average F-measure of 0.742 for overall evaluation, 0.686 for classifying medical problem-treatment relations, 0.820 for medical problem-test relations, and 0.702 for medical problem-medical problem relations. We demonstrate the benefits of learning segment-level representations. We show that medical domain word embeddings help improve relation classification. Seg-CNNs can be trained quickly for the i2b2/VA dataset on a graphics processing unit (GPU) platform. These results support the use of CNNs computed over segments of text for classifying medical relations, as they show state-of-the-art performance while requiring no manual feature engineering. Yuan Luo 0001, Özlem Uzuner, Peter Szolovits, Justin Starren |
J. Am. Medical Informatics Assoc. | 3 |
| 2018 | Corrigendum to "Symptom severity prediction from neuropsychiatric clinical records: Overview of 2016 CEGS N-GRID shared tasks Track 2" [J Biomed Inform. 2017 Nov;75S: S62-S70]
Michele Filannino, Amber Stubbs, Özlem Uzuner |
J. Biomed. Informatics | 3 |
| 2017 | Bridging semantics and syntax with graph algorithms - state-of-the-art of extracting biomedical relationsabstractResearch on extracting biomedical relations has received growing attention recently, with numerous biological and clinical applications including those in pharmacogenomics, clinical trial screening and adverse drug reaction detection. The ability to accurately capture both semantic and syntactic structures in text expressing these relations becomes increasingly critical to enable deep understanding of scientific papers and clinical narratives. Shared task challenges have been organized by both bioinformatics and clinical informatics communities to assess and advance the state-of-the-art research. Significant progress has been made in algorithm development and resource construction. In particular, graph-based approaches bridge semantics and syntax, often achieving the best performance in shared tasks. However, a number of problems at the frontiers of biomedical relation extraction continue to pose interesting challenges and present opportunities for great improvement and fruitful research. In this article, we place biomedical relation extraction against the backdrop of its versatile applications, present a gentle introduction to its general pipeline and shared resources, review the current state-of-the-art in methodology advancement, discuss limitations and point out several promising future directions. Yuan Luo 0001, Özlem Uzuner, Peter Szolovits |
Briefings Bioinform. | 2 |
| 2017 | Bridging semantics and syntax with graph algorithms - state-of-the-art of extracting biomedical relationsabstractBriefings in Bioinformatics (2017) 18(1), 2017, 160–178, doi: 10.1093/bib/bbw001 In the above article, the sentence ‘Wang et al. [84] used Latent Dirichl et al. location to create a semantic representation of biomedical named entities and used Kullback-Leibler (KL) divergence to calculate the association distance between pairs of entities in the Chem2Bio2RDF [149] semantic network’ has been corrected to ‘Wang et al. [84] used Latent Dirichlet Allocation to create a semantic representation of biomedical named entities and used Kullback-Leibler (KL) divergence to calculate the association distance between pairs of entities in the Chem2Bio2RDF [149] semantic network’. The text has been corrected online. The publisher apologizes for this error. Yuan Luo 0001, Özlem Uzuner, Peter Szolovits |
Briefings Bioinform. | 2 |
| 2017 | De-identification of patient notes with recurrent neural networksabstractOBJECTIVE: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the confidentiality of patients. In the United States, the Health Insurance Portability and Accountability Act (HIPAA) defines 18 types of protected health information that needs to be removed to de-identify patient notes. Manual de-identification is impractical given the size of electronic health record databases, the limited number of researchers with access to non-de-identified notes, and the frequent mistakes of human annotators. A reliable automated de-identification system would consequently be of high value. MATERIALS AND METHODS: We introduce the first de-identification system based on artificial neural networks (ANNs), which requires no handcrafted features or rules, unlike existing systems. We compare the performance of the system with state-of-the-art systems on two datasets: the i2b2 2014 de-identification challenge dataset, which is the largest publicly available de-identification dataset, and the MIMIC de-identification dataset, which we assembled and is twice as large as the i2b2 2014 dataset. RESULTS: Our ANN model outperforms the state-of-the-art systems. It yields an F1-score of 97.85 on the i2b2 2014 dataset, with a recall of 97.38 and a precision of 98.32, and an F1-score of 99.23 on the MIMIC de-identification dataset, with a recall of 99.25 and a precision of 99.21. CONCLUSION: Our findings support the use of ANNs for de-identification of patient notes, as they show better performance than previously published systems while requiring no manual feature engineering. Franck Dernoncourt, Ji Young Lee 0001, Özlem Uzuner, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 3 |
| 2017 | Automatic prediction of coronary artery disease from clinical narrativesabstractCoronary Artery Disease (CAD) is not only the most common form of heart disease, but also the leading cause of death in both men and women (Coronary Artery Disease: MedlinePlus, 2015). We present a system that is able to automatically predict whether patients develop coronary artery disease based on their narrative medical histories, i.e., clinical free text. Although the free text in medical records has been used in several studies for identifying risk factors of coronary artery disease, to the best of our knowledge our work marks the first attempt at automatically predicting development of CAD. We tackle this task on a small corpus of diabetic patients. The size of this corpus makes it important to limit the number of features in order to avoid overfitting. We propose an ontology-guided approach to feature extraction, and compare it with two classic feature selection techniques. Our system achieves state-of-the-art performance of 77.4% F1 score. Kevin Buchan, Michele Filannino, Özlem Uzuner |
J. Biomed. Informatics | 3 |
| 2017 | Prescription extraction using CRFs and word embeddingsabstractIn medical practices, doctors detail patients' care plan via discharge summaries written in the form of unstructured free texts, which among the others contain medication names and prescription information. Extracting prescriptions from discharge summaries is challenging due to the way these documents are written. Handwritten rules and medical gazetteers have proven to be useful for this purpose but come with limitations on performance, scalability, and generalizability. We instead present a machine learning approach to extract and organize medication names and prescription information into individual entries. Our approach utilizes word embeddings and tackles the task in two extraction steps, both of which are treated as sequence labeling problems. When evaluated on the 2009 i2b2 Challenge official benchmark set, the proposed approach achieves a horizontal phrase-level F1-measure of 0.864, which to the best of our knowledge represents an improvement over the current state-of-the-art. Carson Tao, Michele Filannino, Özlem Uzuner |
J. Biomed. Informatics | 3 |
| 2016 | Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001 |
AMIA | 11 |
| 2015 | State of the Art of Clinical Narrative Report De-Identification and Its Future
Özlem Uzuner, John S. Aberdeen, Stéphane M. Meystre, Mehmet Kayaalp 0002 |
AMIA | 1 |
| 2015 | Subgraph augmented non-negative tensor factorization (SANTF) for modeling clinical narrative textabstractOBJECTIVE: Extracting medical knowledge from electronic medical records requires automated approaches to combat scalability limitations and selection biases. However, existing machine learning approaches are often regarded by clinicians as black boxes. Moreover, training data for these automated approaches at often sparsely annotated at best. The authors target unsupervised learning for modeling clinical narrative text, aiming at improving both accuracy and interpretability. METHODS: The authors introduce a novel framework named subgraph augmented non-negative tensor factorization (SANTF). In addition to relying on atomic features (e.g., words in clinical narrative text), SANTF automatically mines higher-order features (e.g., relations of lymphoid cells expressing antigens) from clinical narrative text by converting sentences into a graph representation and identifying important subgraphs. The authors compose a tensor using patients, higher-order features, and atomic features as its respective modes. We then apply non-negative tensor factorization to cluster patients, and simultaneously identify latent groups of higher-order features that link to patient clusters, as in clinical guidelines where a panel of immunophenotypic features and laboratory results are used to specify diagnostic criteria. RESULTS AND CONCLUSION: SANTF demonstrated over 10% improvement in averaged F-measure on patient clustering compared to widely used non-negative matrix factorization (NMF) and k-means clustering methods. Multiple baselines were established by modeling patient data using patient-by-features matrices with different feature configurations and then performing NMF or k-means to cluster patients. Feature analysis identified latent groups of higher-order features that lead to medical insights. We also found that the latent groups of atomic features help to better correlate the latent groups of higher-order features. Yuan Luo 0001, Ephraim P. Hochberg, Rohit Joshi, Özlem Uzuner, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Normalization of relative and incomplete temporal expressions in clinical narrativesabstractOBJECTIVE: To improve the normalization of relative and incomplete temporal expressions (RI-TIMEXes) in clinical narratives. METHODS: We analyzed the RI-TIMEXes in temporally annotated corpora and propose two hypotheses regarding the normalization of RI-TIMEXes in the clinical narrative domain: the anchor point hypothesis and the anchor relation hypothesis. We annotated the RI-TIMEXes in three corpora to study the characteristics of RI-TMEXes in different domains. This informed the design of our RI-TIMEX normalization system for the clinical domain, which consists of an anchor point classifier, an anchor relation classifier, and a rule-based RI-TIMEX text span parser. We experimented with different feature sets and performed an error analysis for each system component. RESULTS: The annotation confirmed the hypotheses that we can simplify the RI-TIMEXes normalization task using two multi-label classifiers. Our system achieves anchor point classification, anchor relation classification, and rule-based parsing accuracy of 74.68%, 87.71%, and 57.2% (82.09% under relaxed matching criteria), respectively, on the held-out test set of the 2012 i2b2 temporal relation challenge. DISCUSSION: Experiments with feature sets reveal some interesting findings, such as: the verbal tense feature does not inform the anchor relation classification in clinical narratives as much as the tokens near the RI-TIMEX. Error analysis showed that underrepresented anchor point and anchor relation classes are difficult to detect. CONCLUSIONS: We formulate the RI-TIMEX normalization problem as a pair of multi-label classification problems. Considering only RI-TIMEX extraction and normalization, the system achieves statistically significant improvement over the RI-TIMEX results of the best systems in the 2012 i2b2 challenge. Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 3 |
| 2015 | A systematic comparison of feature space effects on disease classifier performance for phenotype identification of five diseasesabstractAutomated phenotype identification plays a critical role in cohort selection and bioinformatics data mining. Natural Language Processing (NLP)-informed classification techniques can robustly identify phenotypes in unstructured medical notes. In this paper, we systematically assess the effect of naive, lexically normalized, and semantic feature spaces on classifier performance for obesity, atherosclerotic cardiovascular disease (CAD), hyperlipidemia, hypertension, and diabetes. We train support vector machines (SVMs) using individual feature spaces as well as combinations of these feature spaces on two small training corpora (730 and 790 documents) and a combined (1520 documents) training corpus. We assess the importance of feature spaces and training data size on SVM model performance. We show that inclusion of semantically-informed features does not statistically improve performance for these models. The addition of training data has weak effects of mixed statistical significance across disease classes suggesting larger corpora are not necessary to achieve relatively high performance with these models. Christopher Kotfila, Özlem Uzuner |
J. Biomed. Informatics | 2 |
| 2015 | Creation of a new longitudinal corpus of clinical narrativesabstractThe 2014 i2b2/UTHealth Natural Language Processing (NLP) shared task featured a new longitudinal corpus of 1304 records representing 296 diabetic patients. The corpus contains three cohorts: patients who have a diagnosis of coronary artery disease (CAD) in their first record, and continue to have it in subsequent records; patients who do not have a diagnosis of CAD in the first record, but develop it by the last record; patients who do not have a diagnosis of CAD in any record. This paper details the process used to select records for this corpus and provides an overview of novel research uses for this corpus. This corpus is the only annotated corpus of longitudinal clinical narratives currently available for research to the general research community. Vishesh Kumar, Amber Stubbs, Stanley Y. Shaw, Özlem Uzuner |
J. Biomed. Informatics | 4 |
| 2015 | Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth shared task Track 1abstractThe 2014 i2b2/UTHealth Natural Language Processing (NLP) shared task featured four tracks. The first of these was the de-identification track focused on identifying protected health information (PHI) in longitudinal clinical narratives. The longitudinal nature of clinical narratives calls particular attention to details of information that, while benign on their own in separate records, can lead to identification of patients in combination in longitudinal records. Accordingly, the 2014 de-identification track addressed a broader set of entities and PHI than covered by the Health Insurance Portability and Accountability Act - the focus of the de-identification shared task that was organized in 2006. Ten teams tackled the 2014 de-identification task and submitted 22 system outputs for evaluation. Each team was evaluated on their best performing system output. Three of the 10 systems achieved F1 scores over .90, and seven of the top 10 scored over .75. The most successful systems combined conditional random fields and hand-written rules. Our findings indicate that automated systems can be very effective for this task, but that de-identification is not yet a solved problem. Amber Stubbs, Christopher Kotfila, Özlem Uzuner |
J. Biomed. Informatics | 3 |
| 2015 | Identifying risk factors for heart disease over time: Overview of 2014 i2b2/UTHealth shared task Track 2abstractThe second track of the 2014 i2b2/UTHealth natural language processing shared task focused on identifying medical risk factors related to Coronary Artery Disease (CAD) in the narratives of longitudinal medical records of diabetic patients. The risk factors included hypertension, hyperlipidemia, obesity, smoking status, and family history, as well as diabetes and CAD, and indicators that suggest the presence of those diseases. In addition to identifying the risk factors, this track of the 2014 i2b2/UTHealth shared task studied the presence and progression of the risk factors in longitudinal medical records. Twenty teams participated in this track, and submitted 49 system runs for evaluation. Six of the top 10 teams achieved F1 scores over 0.90, and all 10 scored over 0.87. The most successful system used a combination of additional annotations, external lexicons, hand-written rules and Support Vector Machines. The results of this track indicate that identification of risk factors and their progression over time is well within the reach of automated systems. Amber Stubbs, Christopher Kotfila, Hua Xu 0001, Özlem Uzuner |
J. Biomed. Informatics | 4 |
| 2015 | Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpusabstractThe 2014 i2b2/UTHealth natural language processing shared task featured a track focused on the de-identification of longitudinal medical records. For this track, we de-identified a set of 1304 longitudinal medical records describing 296 patients. This corpus was de-identified under a broad interpretation of the HIPAA guidelines using double-annotation followed by arbitration, rounds of sanity checking, and proof reading. The average token-based F1 measure for the annotators compared to the gold standard was 0.927. The resulting annotations were used both to de-identify the data and to set the gold standard for the de-identification track of the 2014 i2b2/UTHealth shared task. All annotated private health information were replaced with realistic surrogates automatically and then read over and corrected manually. The resulting corpus is the first of its kind made available for de-identification research. This corpus was first used for the 2014 i2b2/UTHealth shared task, during which the systems achieved a mean F-measure of 0.872 and a maximum F-measure of 0.964 using entity-based micro-averaged evaluations. Amber Stubbs, Özlem Uzuner |
J. Biomed. Informatics | 2 |
| 2015 | Annotating risk factors for heart disease in clinical narratives for diabetic patientsabstractThe 2014 i2b2/UTHealth natural language processing shared task featured a track focused on identifying risk factors for heart disease (specifically, Cardiac Artery Disease) in clinical narratives. For this track, we used a "light" annotation paradigm to annotate a set of 1304 longitudinal medical records describing 296 patients for risk factors and the times they were present. We designed the annotation task for this track with the goal of balancing annotation load and time with quality, so as to generate a gold standard corpus that can benefit a clinically-relevant task. We applied light annotation procedures and determined the gold standard using majority voting. On average, the agreement of annotators with the gold standard was above 0.95, indicating high reliability. The resulting document-level annotations generated for each record in each longitudinal EMR in this corpus provide information that can support studies of progression of heart disease risk factors in the included patients over time. These annotations were used in the Risk Factor track of the 2014 i2b2/UTHealth shared task. Participating systems achieved a mean micro-averaged F1 measure of 0.815 and a maximum F1 measure of 0.928 for identifying these risk factors in patient records. Amber Stubbs, Özlem Uzuner |
J. Biomed. Informatics | 2 |
| 2015 | Practical applications for natural language processing in clinical research: The 2014 i2b2/UTHealth shared tasksabstract• Capstone shared task for 8 years of i2b2 challenges. Co-organized with UTHealth. • Four tracks: de-identification, risk factor extraction, software usability, and novel data use. • Participation from around the world, from academia and industry. • Data sets available for research beyond the lifetime of i2b2, at i2b2.org/NLP. Özlem Uzuner, Amber Stubbs |
J. Biomed. Informatics | 1 |
| 2015 | Ease of adoption of clinical natural language processing software: An evaluation of five systemsabstractOBJECTIVE: In recognition of potential barriers that may inhibit the widespread adoption of biomedical software, the 2014 i2b2 Challenge introduced a special track, Track 3 - Software Usability Assessment, in order to develop a better understanding of the adoption issues that might be associated with the state-of-the-art clinical NLP systems. This paper reports the ease of adoption assessment methods we developed for this track, and the results of evaluating five clinical NLP system submissions. MATERIALS AND METHODS: A team of human evaluators performed a series of scripted adoptability test tasks with each of the participating systems. The evaluation team consisted of four "expert evaluators" with training in computer science, and eight "end user evaluators" with mixed backgrounds in medicine, nursing, pharmacy, and health informatics. We assessed how easy it is to adopt the submitted systems along the following three dimensions: communication effectiveness (i.e., how effective a system is in communicating its designed objectives to intended audience), effort required to install, and effort required to use. We used a formal software usability testing tool, TURF, to record the evaluators' interactions with the systems and 'think-aloud' data revealing their thought processes when installing and using the systems and when resolving unexpected issues. RESULTS: Overall, the ease of adoption ratings that the five systems received are unsatisfactory. Installation of some of the systems proved to be rather difficult, and some systems failed to adequately communicate their designed objectives to intended adopters. Further, the average ratings provided by the end user evaluators on ease of use and ease of interpreting output are -0.35 and -0.53, respectively, indicating that this group of users generally deemed the systems extremely difficult to work with. While the ratings provided by the expert evaluators are higher, 0.6 and 0.45, respectively, these ratings are still low indicating that they also experienced considerable struggles. DISCUSSION: The results of the Track 3 evaluation show that the adoptability of the five participating clinical NLP systems has a great margin for improvement. Remedy strategies suggested by the evaluators included (1) more detailed and operation system specific use instructions; (2) provision of more pertinent onscreen feedback for easier diagnosis of problems; (3) including screen walk-throughs in use instructions so users know what to expect and what might have gone wrong; (4) avoiding jargon and acronyms in materials intended for end users; and (5) packaging prerequisites required within software distributions so that prospective adopters of the software do not have to obtain each of the third-party components on their own. Kai Zheng 0002, V. G. Vinod Vydiswaran, Yang Liu 0019, Yue Wang 0035, Amber Stubbs, Özlem Uzuner, Anupama E. Gururaj, Samuel Bayer, John S. Aberdeen, Anna Rumshisky, Serguei V. S. Pakhomov, Hua Xu 0001 |
J. Biomed. Informatics | 6 |
| 2014 | Research and applications: Word sense disambiguation in the clinical domain: a comparison of knowledge-rich and knowledge-poor unsupervised methodsabstractOBJECTIVE: To evaluate state-of-the-art unsupervised methods on the word sense disambiguation (WSD) task in the clinical domain. In particular, to compare graph-based approaches relying on a clinical knowledge base with bottom-up topic-modeling-based approaches. We investigate several enhancements to the topic-modeling techniques that use domain-specific knowledge sources. MATERIALS AND METHODS: The graph-based methods use variations of PageRank and distance-based similarity metrics, operating over the Unified Medical Language System (UMLS). Topic-modeling methods use unlabeled data from the Multiparameter Intelligent Monitoring in Intensive Care (MIMIC II) database to derive models for each ambiguous word. We investigate the impact of using different linguistic features for topic models, including UMLS-based and syntactic features. We use a sense-tagged clinical dataset from the Mayo Clinic for evaluation. RESULTS: The topic-modeling methods achieve 66.9% accuracy on a subset of the Mayo Clinic's data, while the graph-based methods only reach the 40-50% range, with a most-frequent-sense baseline of 56.5%. Features derived from the UMLS semantic type and concept hierarchies do not produce a gain over bag-of-words features in the topic models, but identifying phrases from UMLS and using syntax does help. DISCUSSION: Although topic models outperform graph-based methods, semantic features derived from the UMLS prove too noisy to improve performance beyond bag-of-words. CONCLUSIONS: Topic modeling for WSD provides superior results in the clinical domain; however, integration of knowledge remains to be effectively exploited. Rachel Chasin, Anna Rumshisky, Özlem Uzuner, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | Evaluating temporal relations in clinical text: 2012 i2b2 ChallengeabstractBACKGROUND: The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives. The organizers provided the research community with a corpus of discharge summaries annotated with temporal information, to be used for the development and evaluation of temporal reasoning systems. 18 teams from around the world participated in the challenge. During the workshop, participating teams presented comprehensive reviews and analysis of their systems, and outlined future research directions suggested by the challenge contributions. METHODS: The challenge evaluated systems on the information extraction tasks that targeted: (1) clinically significant events, including both clinical concepts such as problems, tests, treatments, and clinical departments, and events relevant to the patient's clinical timeline, such as admissions, transfers between departments, etc; (2) temporal expressions, referring to the dates, times, durations, or frequencies phrases in the clinical text. The values of the extracted temporal expressions had to be normalized to an ISO specification standard; and (3) temporal relations, between the clinical events and temporal expressions. Participants determined pairs of events and temporal expressions that exhibited a temporal relation, and identified the temporal relation between them. RESULTS: For event detection, statistical machine learning (ML) methods consistently showed superior performance. While ML and rule based methods seemed to detect temporal expressions equally well, the best systems overwhelmingly adopted a rule based approach for value normalization. For temporal relation classification, the systems using hybrid approaches that combined ML and heuristics based methods produced the best results. Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | Temporal reasoning over clinical text: the state of the artabstractOBJECTIVES: To provide an overview of the problem of temporal reasoning over clinical text and to summarize the state of the art in clinical natural language processing for this task. TARGET AUDIENCE: This overview targets medical informatics researchers who are unfamiliar with the problems and applications of temporal reasoning over clinical text. SCOPE: We review the major applications of text-based temporal reasoning, describe the challenges for software systems handling temporal information in clinical text, and give an overview of the state of the art. Finally, we present some perspectives on future research directions that emerged during the recent community-wide challenge on text-based temporal reasoning in the clinical domain. Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | Annotating temporal information in clinical narratives
Weiyi Sun, Anna Rumshisky, Özlem Uzuner |
J. Biomed. Informatics | 3 |
| 2013 | Chronology of your health events: Approaches to extracting temporal relations from medical narratives
Özlem Uzuner, Amber Stubbs, Weiyi Sun |
J. Biomed. Informatics | 1 |
| 2012 | Using UMLS for Word Sense Disambiguation in Clinical Notes
Anna Rumshisky, Rachel Chasin, Özlem Uzuner, Peter Szolovits |
AMIA | 3 |
| 2012 | Evaluation and Visualization of Human Annotator Learning Patterns
Ying Suo, Shuying Shen, Scott L. DuVall, Özlem Uzuner, Brett R. South |
AMIA | 4 |
| 2012 | MCORES: a system for noun phrase coreference resolution for clinical recordsabstractOBJECTIVE: Narratives of electronic medical records contain information that can be useful for clinical practice and multi-purpose research. This information needs to be put into a structured form before it can be used by automated systems. Coreference resolution is a step in the transformation of narratives into a structured form. METHODS: This study presents a medical coreference resolution system (MCORES) for noun phrases in four frequently used clinical semantic categories: persons, problems, treatments, and tests. MCORES treats coreference resolution as a binary classification task. Given a pair of concepts from a semantic category, it determines coreferent pairs and clusters them into chains. MCORES uses an enhanced set of lexical, syntactic, and semantic features. Some MCORES features measure the distance between various representations of the concepts in a pair and can be asymmetric. RESULTS AND CONCLUSION: MCORES was compared with an in-house baseline that uses only single-perspective 'token overlap' and 'number agreement' features. MCORES was shown to outperform the baseline; its enhanced features contribute significantly to performance. In addition to the baseline, MCORES was compared against two available third-party, open-domain systems, RECONCILE(ACL09) and the Beautiful Anaphora Resolution Toolkit (BART). MCORES was shown to outperform both of these systems on clinical records. Andreea Bodnari, Peter Szolovits, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Evaluating the state of the art in coreference resolution for electronic medical recordsabstractBACKGROUND: The fifth i2b2/VA Workshop on Natural Language Processing Challenges for Clinical Records conducted a systematic review on resolution of noun phrase coreference in medical records. Informatics for Integrating Biology and the Bedside (i2b2) and the Veterans Affair (VA) Consortium for Healthcare Informatics Research (CHIR) partnered to organize the coreference challenge. They provided the research community with two corpora of medical records for the development and evaluation of the coreference resolution systems. These corpora contained various record types (ie, discharge summaries, pathology reports) from multiple institutions. METHODS: The coreference challenge provided the community with two annotated ground truth corpora and evaluated systems on coreference resolution in two ways: first, it evaluated systems for their ability to identify mentions of concepts and to link together those mentions. Second, it evaluated the ability of the systems to link together ground truth mentions that refer to the same entity. Twenty teams representing 29 organizations and nine countries participated in the coreference challenge. RESULTS: The teams' system submissions showed that machine-learning and rule-based approaches worked best when augmented with external knowledge sources and coreference clues extracted from document structure. The systems performed better in coreference resolution when provided with ground truth mentions. Overall, the systems struggled in solving coreference resolution for cases that required domain knowledge. Özlem Uzuner, Andreea Bodnari, Shuying Shen, Tyler Forbush, John Pestian, Brett R. South |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutionsabstractThis issue of JAMIA focuses on natural language processing (NLP) techniques for clinical-text information extraction. Several articles are offshoots of the yearly ‘Informatics for Integrating Biology and the Bedside’ (i2b2) (http://www.i2b2.org) NLP shared-task challenge, introduced by Uzuner et al (see page 552)1 and co-sponsored by the Veteran's Administration for the last 2 years. This shared task follows long-running challenge evaluations in other fields, such as the Message Understanding Conference (MUC) for information extraction,2 TREC3 for text information retrieval, and CASP4 for protein structure prediction. Shared tasks in the clinical domain are recent and include annual i2b2 Challenges that began in 2006, a challenge for multi-label classification of radiology reports sponsored by Cincinnati Children's Hospital in 2007,5 a 2011 Cincinnati Children's Hospital challenge on suicide notes,6 and the 2011 TREC information retrieval shared task involving retrieval of clinical cases from narrative records.7 Although NLP research in the clinical domain has been active since the 1960s, progress in the development of NLP applications for clinical text has been slow and lags behind progress in the general NLP domain. There are several barriers to NLP development in the clinical domain, and shared tasks like the i2b2/VA Challenge address some of these barriers. Nevertheless, many barriers remain and unless the community takes a more active role in developing novel approaches for addressing the barriers, advancement and innovation will continue to be slow. Historically, there have been substantial barriers to NLP development in the clinical domain. These barriers are not unique to the clinical domain: they also occur in the fields of software engineering and general NLP. Because of concerns regarding patient privacy and worry about revealing unfavorable institutional practices, hospitals and clinics have been extremely reluctant to allow access to clinical data for researchers from outside the associated institutions. The lack of reliable and inexpensive de-identification techniques for narrative reports has compounded the reluctance to share. Such restricted access to shared datasets has hindered collaboration and inhibited the ability to assess and adapt NLP technologies across institutions and among research groups. Several pioneering efforts5,8–11 have made clinical data available for sharing—we need more of these grass-roots efforts. Closely related but not completely conditional on lack of shared datasets is the deficiency of annotated clinical data for training NLP applications and benchmarking performance. The sublanguage of clinical reports often necessitates domain-specific development and training, and, as a consequence, NLP modules developed for general text typically do not perform as well on clinical narratives. We need increased coordination to create annotation sets that can be merged to produce larger training and evaluation sets. Without the ability to share data, the community has lacked incentives for developing common data models for manual and automatic annotations. The result is that annotated datasets are usually unique to the laboratory that generated them and thus remain small and that NLP modules that perform the same tasks cannot be substituted and compared without considerable translational effort. At present, the clinical NLP community is leveraging existing standards and conventions and working together to develop shared data models and to map annotations across information extraction applications. Adopting an existing NLP application or module is complicated—source code and documentation may be unavailable, and published descriptions may lack sufficient detail for reproducibility. Open source releases of clinical information extraction and retrieval systems have improved the opportunity to reproduce performance.12–15 Even with open source release, a tool may work less well in others' hands than in the hands of the original developers. Compounding the problem of reproducibility is the fact that proof-of-concept tools created in academic/research environments may not meet the highest software engineering quality, maintainability, scalability, or usability standards. And sometimes a tool may be over-fitted to a particular application, and modification to solve a similar problem may require wholesale changes. As Pedersen asserted,16 the NLP community needs to invest more in assisting others in applying and reproducing our results. In part due to previously listed barriers, collaboration within the clinical NLP community has been nominal. Development of NLP systems within the academic environment has centered around single institutions and single laboratories, and rather than building upon the foundations of previous work, the majority of clinical NLP systems developed over the last four decades have been reinvented as silos that are neither expanded nor applied outside of the individual laboratory. Other factors limiting collaboration include insufficient infrastructure for facilitating cooperation and the reality that collaboration is inherently inefficient. Nevertheless, as with the biomedical research community at large, a surge in progression beyond the last half century of research can only come through enhanced teamwork. A recent trend for teamwork across NLP research laboratories is evident in funded initiatives such as the VA's Consortium for Healthcare Informatics Research (CHIR)17 and in the ONC-funded SHARP Area 4 grant for Secondary Use of EHR data.18 Also, advances in open source development and conformity to common frameworks has led to recent advances in NLP allowing one team to extend work by another (eg, HiTEX12 built on GATE and cTAKES,14 ODIE,19 and Automated Retrieval Console (ARC) built on UIMA16). Although we are improving incrementally the predictive performance of clinical NLP tools, clinical NLP applications are seldom deployed in clinical, public health, or health services research settings. Currently, the perceived cost of applying NLP outweighs the perceived benefit. Deploying an NLP system typically requires a substantial amount of time from an expert NLP developer—normally, applications do not generalize and must be rebuilt, retrained, enhanced, and re-evaluated for each new task; the output of an NLP system typically requires extensive mapping to the specific problem being addressed; and the ability to aid a user in customizing the application is generally inadequate (see page 544).20 We need a shift of focus from accuracy in one task to generalizabiliy across many and from the production of papers as the sole output to production of usable software for medically relevant applications. We also need to understand where NLP tools fit into an overall user workflow so that the tools can be integrated into end-to-end applications for clinical, public health, and clinical research users. Shared tasks like the i2b2/VA Challenge address several of these barriers in part. Shared tasks provide annotated datasets to participants and sometimes to non-participants (i2b2 datasets are available to others a year after the Challenge). The i2b2 shared task is standardizing its corpus as much as possible—the same records are used from one year to the next with layers of annotation that build on each other, and common input/output specifications are applied every year. Shared tasks partially address the barrier of reproducibility by providing an evaluation opportunity that minimizes the risk of over-fitting: participants have time to train their systems in supervised fashion with an annotated training dataset, but then evaluation must be performed against a separate non-annotated dataset within a stringent time limit that prevents non-trivial system modifications. Although shared tasks are not designed for this purpose, the i2b2 Challenge has been the impetus for some new collaborations across independent research groups. Shared tasks have driven progress in related fields. For example, progress in speech understanding research was driven by a series of evaluations funded by DARPA from the late 1980s to the early 2000s.21 The research community was able to consistently drive down the error rate by a factor of two every 2 years, on successively more challenging tasks, moving from recognition of small-vocabulary read speech to automated transcription of broadcast news in multiple languages. Associated with this progress was the incorporation of speech recognition products into applications, from dictation to speech interfaces. Shared tasks provide value to the NLP community in several ways: Common evaluation metrics are developed. Annotated datasets are made available. Enticed by available annotated datasets, researchers in overlapping fields (both academic and corporate) participate in the tasks, bringing in new people and new approaches. Benchmarking evaluation on a shared dataset reveals the state-of-the-art performance for a given task. Students and post-docs receive excellent training opportunities. Preliminary results can be obtained by a new research group, which can potentially lead to funding opportunities. Pre-processed, standardized corpora with multiple layers of annotations on the same corpus pave the way for end-to-end evaluations in addition to evaluation on a single annotation layer. Conventions for standardizing annotations and input/output formats are developed, and despite other standardization efforts, shared task corpora often set de facto standards. In spite of the value of shared tasks, the tasks have several shortcomings: Participants come mainly from teams with funded projects that overlap with the shared task. For academic participants, a significant motivation is the opportunity to publish; however, there is sometimes limited value for the larger community in publications resulting from a shared task. Because development time is limited during shared tasks, participants often build on applications that already exist and apply methods already described in the literature. This can result in many similar approaches being applied to the same task. Although publishing the high-performing systems can be interesting, the resulting publications may not be novel and therefore may not improve the general body of knowledge. If a particular challenge task is repeated over time, there is a tendency for system approaches to converge on the approach that showed most success in the previous evaluation—evaluations repeated over time tend to reduce the diversity of approaches. Although shared tasks contribute to growth and progress, increased benefit to the community of clinical NLP developers and to potential users will require additional individual and community efforts that target existing barriers creatively. Driving progress in a way that will increase the impact of NLP in the realm of individual and population health will require creativity at both the grass-roots and the community levels. No single activity can tackle all barriers. In addition to encouraging variations on the development of shared tasks and their incentives, we would like to see new types of shared activities that foster the outcomes described below. In exchange for access to the costly annotated dataset, shared task participation could be contingent on depositing code in a shared repository or creating a web service for prospective users. In this model, the organizers could send test data to the participants' servers and the servers return the results for evaluation (see Leitner et al22 for a description of a metaserver used in evaluation of results from BioCreative II). The servers (and metaserver) could even persist, providing services to interested users beyond the initial shared task. Publication of computational methods in biomedical informatics journals like JAMIA could further encourage reproducibility of results through policies mandating simultaneous submission of code with a manuscript, as recommended by Pedersen.16 The clinical NLP community could independently accelerate reproducibility (and lead by example) by depositing code and developing web services in a common repository.23 This trend is occurring in settings restricted by affiliation, such as the VA VINCI framework24 (available only to VA researchers) and a cloud environment being hosted by the SHARP Area 4 grant18 (available to grant participants). The new National Center for Biomedical Computing iDASH25 is developing a similar cyber-infrastructure that will be restricted not by affiliation but by adherence to privacy policies and agreements required by data contributors. The National Library of Medicine is currently hosting a registry developed by the AMIA NLP working group called ORBIT for listing and pointing to biomedical informatics and NLP resources.26 In a shared task, dozens of research groups duplicate the same task independently. Although a variety of techniques can emerge for the same task, given the relatively short time frame allowed for development and training, the features and approaches applied in the challenge are often very similar. Whereas similarity of approaches reveals agreement among teams on the best approaches and sets the stage for collaboration, the competitive nature of shared tasks provides a disincentive to collaboration; the reward system for shared tasks is not at all dependent on the ability to collaborate across teams but is solely geared toward competition in which a single winner arises. The open source development community has found inherent rewards in collaborative development through supportive environments like GitHub. Perhaps we can learn from collaborative development communities who participate in hackathons27 and from the games industry in asking how a shared task can be designed so that collaboration is rewarded and becomes worthwhile, interesting, and attractive.28 Evaluation of a shared task is focused on accuracy, and existing challenges evaluate only predictive performance, not software engineering characteristics or usability. Imagine a shared task in which success is judged on usability of a system or direct portability of one technique to a new task or domain. Evaluating success of a system with this paradigm is inherently more complex, but we could learn from groupware evaluation, from the rich field of usability testing, and from incentivized competitions like those sponsored by the X-Prize Foundation.29 Because of the cost of creating annotated training data, shared tasks are often small scale, at least relative to real medical applications. We need new approaches to rapid adaptation of NLP systems to new applications, with less dependence on ‘deeply annotated’ data; such applications would present important opportunities for collaboration with the end user community, who might be motivated to provide domain expertise if they were likely to get a scalable, maintainable system out of the collaboration. Scalability will require more efficient techniques for manual annotation. And scalability will require an enriched ability to produce high quality software, which may necessitate better collaboration with industry30 and funding models that include support for operational development. The shared i2b2 evaluations have made a huge contribution to stimulating and vitalizing the field of clinical NLP; however, to ensure the transition into usable applications, the clinical NLP research community needs to address the critical issues of data access, development of shared infrastructure, and integration of software engineering methods to ensure the usability, maintainability, and availability of clinical NLP tools that are integrated into the workflow of real biomedical applications. This must be done in close collaboration with end users, software engineers, and clinical practitioners. We as a community need to think beyond the status quo of incremental improvement in the F score toward imaginative approaches that encourage collaboration, promote reproducibility, increase the scalability of NLP development, and provide value to end users. Authors are funded in part by U54HL108460, R01GM090187, SHARP ONC award 90TR0002, R01 CA127979, and U54LM008748. None. Not commissioned; internally peer reviewed. Wendy W. Chapman, Prakash M. Nadkarni, Lynette Hirschman, Leonard W. D'Avolio, Guergana K. Savova, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 6 |
| 2011 | 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical textabstractThe 2010 i2b2/VA Workshop on Natural Language Processing Challenges for Clinical Records presented three tasks: a concept extraction task focused on the extraction of medical concepts from patient reports; an assertion classification task focused on assigning assertion types for medical problem concepts; and a relation classification task focused on assigning relation types that hold between medical problems, tests, and treatments. i2b2 and the VA provided an annotated reference standard corpus for the three tasks. Using this reference standard, 22 systems were developed for concept extraction, 21 for assertion classification, and 16 for relation classification. These systems showed that machine learning approaches could be augmented with rule-based systems to determine concepts, assertions, and relations. Depending on the task, the rule-based systems can either provide input for machine learning or post-process the output of machine learning. Ensembles of classifiers, information from unlabeled data, and external knowledge sources can help when the training data are inadequate. Özlem Uzuner, Brett R. South, Shuying Shen, Scott L. DuVall |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Semantic relations for problem-oriented medical records
Özlem Uzuner, Jonathan Mailoa, Russell Ryan, Tawanda C. Sibanda |
Artif. Intell. Medicine | 1 |
| 2010 | Extracting medication information from clinical textabstractThe Third i2b2 Workshop on Natural Language Processing Challenges for Clinical Records focused on the identification of medications, their dosages, modes (routes) of administration, frequencies, durations, and reasons for administration in discharge summaries. This challenge is referred to as the medication challenge. For the medication challenge, i2b2 released detailed annotation guidelines along with a set of annotated discharge summaries. Twenty teams representing 23 organizations and nine countries participated in the medication challenge. The teams produced rule-based, machine learning, and hybrid systems targeted to the task. Although rule-based systems dominated the top 10, the best performing system was a hybrid. Of all medication-related fields, durations and reasons were the most difficult for all systems to detect. While medications themselves were identified with better than 0.75 F-measure by all of the top 10 systems, the best F-measure for durations and reasons were 0.525 and 0.459, respectively. State-of-the-art natural language processing systems go a long way toward extracting medication names, dosages, modes, and frequencies. However, they are limited in recognizing duration and reason fields and would benefit from future research. Özlem Uzuner, Imre Solti, Eithon Cadag |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Community annotation experiment for ground truth generation for the i2b2 medication challengeabstractOBJECTIVE: Within the context of the Third i2b2 Workshop on Natural Language Processing Challenges for Clinical Records, the authors (also referred to as 'the i2b2 medication challenge team' or 'the i2b2 team' for short) organized a community annotation experiment. DESIGN: For this experiment, the authors released annotation guidelines and a small set of annotated discharge summaries. They asked the participants of the Third i2b2 Workshop to annotate 10 discharge summaries per person; each discharge summary was annotated by two annotators from two different teams, and a third annotator from a third team resolved disagreements. MEASUREMENTS: In order to evaluate the reliability of the annotations thus produced, the authors measured community inter-annotator agreement and compared it with the inter-annotator agreement of expert annotators when both the community and the expert annotators generated ground truth based on pooled system outputs. For this purpose, the pool consisted of the three most densely populated automatic annotations of each record. The authors also compared the community inter-annotator agreement with expert inter-annotator agreement when the experts annotated raw records without using the pool. Finally, they measured the quality of the community ground truth by comparing it with the expert ground truth. RESULTS AND CONCLUSIONS: The authors found that the community annotators achieved comparable inter-annotator agreement to expert annotators, regardless of whether the experts annotated from the pool. Furthermore, the ground truth generated by the community obtained F-measures above 0.90 against the ground truth of the experts, indicating the value of the community as a source of high-quality ground truth even on intricate and domain-specific annotation tasks. Özlem Uzuner, Imre Solti, Fei Xia 0004, Eithon Cadag |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Viewpoint Paper: Recognizing Obesity and Comorbidities in Sparse DataabstractIn order to survey, facilitate, and evaluate studies of medical language processing on clinical narratives, i2b2 (Informatics for Integrating Biology to the Bedside) organized its second challenge and workshop. This challenge focused on automatically extracting information on obesity and fifteen of its most common comorbidities from patient discharge summaries. For each patient, obesity and any of the comorbidities could be Present, Absent, or Questionable (i.e., possible) in the patient, or Unmentioned in the discharge summary of the patient. i2b2 provided data for, and invited the development of, automated systems that can classify obesity and its comorbidities into these four classes based on individual discharge summaries. This article refers to obesity and comorbidities as diseases. It refers to the categories Present, Absent, Questionable, and Unmentioned as classes. The task of classifying obesity and its comorbidities is called the Obesity Challenge. The data released by i2b2 was annotated for textual judgments reflecting the explicitly reported information on diseases, and intuitive judgments reflecting medical professionals' reading of the information presented in discharge summaries. There were very few examples of some disease classes in the data. The Obesity Challenge paid particular attention to the performance of systems on these less well-represented classes. A total of 30 teams participated in the Obesity Challenge. Each team was allowed to submit two sets of up to three system runs for evaluation, resulting in a total of 136 submissions. The submissions represented a combination of rule-based and machine learning approaches. Evaluation of system runs shows that the best predictions of textual judgments come from systems that filter the potentially noisy portions of the narratives, project dictionaries of disease names onto the remaining text, apply negation extraction, and process the text through rules. Information on disease-related concepts, such as symptoms and medications, and general medical knowledge help systems infer intuitive judgments on the diseases. Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Research Paper: Machine Learning and Rule-based Approaches to Assertion ClassificationabstractOBJECTIVES: The authors study two approaches to assertion classification. One of these approaches, Extended NegEx (ENegEx), extends the rule-based NegEx algorithm to cover alter-association assertions; the other, Statistical Assertion Classifier (StAC), presents a machine learning solution to assertion classification. DESIGN: For each mention of each medical problem, both approaches determine whether the problem, as asserted by the context of that mention, is present, absent, or uncertain in the patient, or associated with someone other than the patient. The authors use these two systems to (1) extend negation and uncertainty extraction to recognition of alter-association assertions, (2) determine the contribution of lexical and syntactic context to assertion classification, and (3) test if a machine learning approach to assertion classification can be as generally applicable and useful as its rule-based counterparts. MEASUREMENTS: The authors evaluated assertion classification approaches with precision, recall, and F-measure. RESULTS: The ENegEx algorithm is a general algorithm that can be directly applied to new corpora. Despite being based on machine learning, StAC can also be applied out-of-the-box to new corpora and achieve similar generality. CONCLUSION: The StAC models that are developed on discharge summaries can be successfully applied to radiology reports. These models benefit the most from words found in the +/- 4 word window of the target and can outperform ENegEx. Özlem Uzuner, Tawanda C. Sibanda |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Specializing for predicting obesity and its co-morbiditiesabstractWe present specializing, a method for combining classifiers for multi-class classification. Specializing trains one specialist classifier per class and utilizes each specialist to distinguish that class from all others in a one-versus-all manner. It then supplements the specialist classifiers with a catch-all classifier that performs multi-class classification across all classes. We refer to the resulting combined classifier as a specializing classifier. We develop specializing to classify 16 diseases based on discharge summaries. For each discharge summary, we aim to predict whether each disease is present, absent, or questionable in the patient, or unmentioned in the discharge summary. We treat the classification of each disease as an independent multi-class classification task. For each disease, we develop one specialist classifier for each of the present, absent, questionable, and unmentioned classes; we supplement these specialist classifiers with a catch-all classifier that encompasses all of the classes for that disease. We evaluate specializing on each of the 16 diseases and show that it improves significantly over voting and stacking when used for multi-class classification on our data. Ira Goldstein, Özlem Uzuner |
J. Biomed. Informatics | 2 |
| 2008 | Two Approaches to Assertion Classification
Özlem Uzuner, Tawanda C. Sibanda |
AMIA | 1 |
| 2008 | A de-identifier for medical discharge summaries
Özlem Uzuner, Tawanda C. Sibanda, Yuan Luo 0001, Peter Szolovits |
Artif. Intell. Medicine | 1 |
| 2008 | Letters to the Editor: No Structure Before Its TimeabstractDr. Schleyer asks a number of important questions. These might be summarized as asking why, as a discipline, are we not focusing on improving the acquisition of structured data rather than going through computational acrobatics to extract codified representation from narrative text? Should we not be focusing our efforts to ensure a fully-structured record? We agree with Dr. Schleyer that a fully-structured record is an important goal. In the context of a healthcare system with 10-minute visits, time of data entry using current technologies remains burdensome.1 Even if time of data entry were of no concern, the engineering of structured data acquisition interfaces to support the full richness of patient state captured by natural language remains a challenge. Most tellingly, there have been over four decades of research by leading informaticians in precisely the area that Dr. Schleyer urges us to explore and yet the volume of narrative text continues to grow in our leading academic centers. Perhaps the greatest potential for developing a fully-structured medical record lies in patients' annotating their own record given their direct interest in accuracy and disposable time. Forty years ago, it was shown that patients could accurately present their symptoms in codified fashion2 and subsequent research suggests that they and their families can do equally well in reporting medications and even physiological state.3,4 With the increased popularity of personally-controlled health records, such capabilities are likely to be increasingly exploited. Notwithstanding these hopeful developments, it seems quite likely that the amount of narrative text describing patient states will continue to grow in the near to mid-term and the existing corpus will persist for at least the life of our patients. The obligation and the opportunity to provide the best possible care of our patients and the most informed research using such data will therefore continue to motivate segments of the academic and commercial informatics research communities in refining natural language processing techniques. Isaac S. Kohane, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 2 |
| 2008 | Viewpoint Paper: Identifying Patient Smoking Status from Medical Discharge RecordsabstractClinical narrative records contain much useful information. However, most clinical narratives are in the form of fragmented English free text, showing the characteristics of a clinical sublanguage. This makes their linguistic processing, search, and retrieval challenging.1 Traditional natural language processing (NLP) tools are not designed for the fragmented free text found in narrative clinical records; therefore, they do not perform well on this type of data.2 Limited access to clinical records has been a barrier to the widespread development of medical language processing (MLP) technologies. In the absence of a standardized, publicly available ground truth that encourages the development of MLP systems and allows their head-to-head comparison, successful MLP efforts have been limited, e.g., MedLEE3 and Symtxt.4 A few MLP systems have been developed,5 and such efforts have successfully shown the usefulness of MLP in clinical settings.6–8 To improve the availability of clinical records and to contribute to the advancement of the state of the art in MLP, within the i2b2 (Informatics for Integrating Biology to the Bedside) project, the authors de-identified and released a set of clinical records from Partners HealthCare. These records provided the basis for the development of ground truth for two challenge questions: Automatic de-identification of clinical data, i.e., de-identification challenge. Automatic evaluation of the smoking status of patients based on medical records, i.e., smoking challenge. Representative teams from the MLP community participated in the two challenges and met at a workshop organized by the authors to discuss the results of the challenges. The workshop was co-sponsored by the American Medical Informatics Association and met in conjunction with its Fall Symposium in November 2006. This article provides an overview of the smoking challenge and the findings of the workshop. An overview of the de-identification challenge can be found in Uzuner et al.9 The smoking challenge continues the tradition of attempting to identify the state of the art in automatic language processing. Outside of the medical domain, there have been many efforts in this direction. Most of these efforts have been led by Message Understanding Conferences (MUC)10 and the National Institute of Standards and Technology (NIST).11 MUC organized shared tasks on named entity recognition. NIST organized a series of Text Retrieval Evaluation Conferences (TREC) on various domains including blogs and legal documents; they also organized a series of shared tasks on topic detection and tracking, speaker recognition, language recognition, spoken document retrieval, machine translation, and entity extraction. In the biomedical domain, three such prominent efforts were BioCreAtIvE12 for information extraction, ImageCLEF13,14 for image retrieval, and TREC Genomics15 for question answering and information retrieval. Keeping the goals of TREC,16 MUC,17 BioCreAtIvE,18 etc., in mind for the smoking challenge, we created a collection of actual medical discharge records. We invited the development of systems that can predict the smoking status of patients based on the narratives in these medical discharge records. We limited the scope of this task to understanding only the explicitly reported smoking information. In other words, information that implicitly reveals the smoking status was excluded from this study. Our smoking challenge continued the work on the application of classification techniques to the medical domain,19–22 and extended the MLP studies on medical discharge records.8,23–29 Information on the smoking status of patients is important for many health studies, e.g., studies on asthma; however, before this challenge, the only system for the automatic evaluation of the smoking status of patients from their records was the HITEx system.30 The data for the smoking challenge consisted exclusively of discharge summaries from Partners HealthCare. We preprocessed these records so that they were de-identified, tokenized, broken into sentences, converted into XML format, and separated into training and test sets. Institutional review boards of Partners HealthCare, Massachusetts Institute of Technology, and the State University of New York at Albany approved the challenge and the data preparation process. The data for the challenge were annotated by pulmonologists. The pulmonologists were asked to classify patient records into five possible smoking status categories. For the purposes of this challenge, we defined these categories as follows: A Past Smoker is a patient whose discharge summary asserts explicitly that the patient was a smoker one year or more ago but who has not smoked for at least one year. The assertion “past smoker” without any temporal qualifications means Past Smoker unless there is text that says that the patient stopped smoking less than one year ago. A Current Smoker is a patient whose discharge summary asserts explicitly that the patient was a smoker within the past year. The assertion “current smoker” without any temporal qualifications means Current Smoker unless there is text that says that the patient stopped smoking more than a year ago. A Smoker is a patient who is either a Current or a Past Smoker but whose medical record does not provide enough information to classify the patient as either. A Non-Smoker's discharge summary indicates that they never smoked. An Unknown is a patient whose discharge summary does not mention anything about smoking. Indecision between Current Smoker and Past Smoker does not belong to this category. Second-hand smokers are considered Non-Smokers for the purposes of this study, unless there is evidence in their record that they actively smoked. Similarly, as we are only concerned with tobacco, marijuana smoking should not affect the patients' smoking status. In addition to being provided with the above definitions, the annotators were trained on 55 sample sentences30 (see Table 1 for a subset) and 10 sample records. Annotator Training Samples Annotator Training Samples Two pulmonologists annotated each record with the smoking status of patients based strictly on the explicitly stated smoking-related facts in the records. These annotations constitute the textual judgments of the annotators. The same two pulmonologists also marked the smoking status of the patients using their medical intuitions on all information in the records. These annotations constitute the intuitive judgments of the annotators. In all, 928 records were annotated. The interannotator agreement on the textual judgments on these records, as measured by Cohen's kappa (κ),31,32 was 0.84; observed agreement on textual judgments was 0.93; specific agreement per category on textual judgments ranged from 0.4 to 0.98 (Table 2; also see the Methods section for definitions of Cohen's kappa, observed agreement, and specific agreement). The interannotator agreement on the intuitive judgments was 0.45; observed agreement on intuitive judgments was 0.73; specific agreement per category on intuitive judgments ranged from 0.3 to 0.84 (Table 2). Observed and Specific Agreement Observed and Specific Agreement Guidelines for κ are subject to interpretation and depend on parameters such as the task and categories involved.33 However, κ of 0.8 is widely used as the threshold for strong agreement.31–35 On our data, we observed strong agreement only on the textual judgments. We further observed that the intra-annotator agreement between a given doctor's intuitive and textual judgments varied from 0.62 to 0.99. This indicates that the reliance of the intuitive judgments on the explicit textual information varies considerably from doctor to doctor. Given these observations, we limited the challenge task to the identification of the smoking status based on information that is explicitly mentioned in the records. To generate the ground truth, we resolved the disagreements (on textual judgments) between the annotators by obtaining judgments from two other pulmonologists. We omitted the records that the annotators disagreed on from the challenge, unless a majority vote could identify a clear textual judgment for them. In all, 63 records were omitted from the challenge for lack of a clear textual judgment. In addition, annotation results showed heavy bias in the data for Unknown records. This is the least interesting category for the purposes of the smoking challenge as Unknown records do not contain any smoking-related information. To focus the smoking challenge less on the Unknown category and more on the other four categories, we omitted a portion of the Unknown records (363 records) from the challenge. A total of 502 de-identified medical discharge records were used for the smoking challenge. Table 3 shows the distribution of annotated records into training and test sets, and into Past Smoker, Current Smoker, Smoker, Non-Smoker, and Unknown categories. The training and test sets show similar distribution of records into the five smoking categories; however, these distributions are far from uniform. This reflects the realities of real-world data; our records were drawn at random from the Partners' database, in which some smoking categories are better represented than others. In our test set, the smallest smoking category is Smokers, with only three records. The training and test data can be obtained from i2b2.org. Smoking Status Training and Test Data Distribution Smoking Status Training and Test Data Distribution We evaluated system performances using microaveraged and macroaveraged precision, recall, and F-measure, as well as Cohen's kappa. Precision, recall, and F-measure are performance metrics frequently used in NLP.36,37 These metrics are easily derived from a binary confusion matrix. In a binary decision problem, a classifier labels entities as either positive or negative (where positive and negative represent two generic categories) and produces a confusion matrix. This matrix contains four entities: true positive (TP), true negative (TN), false positive (FP), and false negative (FN). Given such a matrix, precision is the percentage of entities classified correctly to be in a given category in relation to the total number of entities classified for the given category (Equation 1). Recall is the percentage of entities classified correctly in a given category in relation to the actual number of items in the given category (Equation 2). F-measure is the harmonic mean of precision and recall (Equation 3). β enables F-measure to favor either precision or recall. We give equal weight to precision and recall by setting β = 1. Cohen's kappa (κ) (Equation 6) is a measure of agreement31 between pairs of annotators who classify items into a set number of mutually exclusive categories. κ depends on observed agreement (Ao in Equation 7) and the agreement expected due to chance (Ae in Equation 8). A κ value of 0.8 is widely used as the threshold for strong agreement,31–35 whereas a κ of 0 indicates that the observed agreement is due to chance.33 We used κ as a measure of inter-annotator and intra-annotator agreement (see Annotations section) as a measure of agreement between two automatic systems and as a measure of agreement between a system and the ground truth (see Results and Discussion). Equation 6 through Equation 8 collectively describe κ between an automatic system and the ground truth. κ for inter-annotator and intra-annotator agreement and for inter-system agreement can be computed analogously. According to Hripcsak and Rothschild,38 there exists a correspondence between F-measure and κ. However, κ provides clearer insights into the relative strengths of the systems (see Intersystem Agreement section). We evaluated systems using F-measure but compared them using both F-measure and κ. Specific agreement (Asp)33 measures the degree of agreement (Equation 9) on each category and is not adjusted by chance. We used specific agreement to get a sense of the level of agreement between annotators without taking chance into consideration. We tested the significance of the differences of the systems using a randomization technique that is frequently utilized in NLP.39 The null hypothesis is that the absolute value of the difference in performances, e.g., F-measures, of two systems is approximately equal to zero. The randomization technique does not assume a particular distribution of the differences. Instead, it empirically generates the distribution. Given two actual systems, it randomly shuffles (at each iteration, we simulated a coin flip to decide whether the answers should be swapped) their responses to the records in the test set N times (e.g., N = 9,999), and thus creates N pairs of pseudosystems. It counts the number of times that the difference between the performances of pairs of pseudosystems is greater than the difference between the two actual systems' performances. Let this count be equal to n and compute . If s is greater than a predetermined cutoff α, then the difference of the performances of the two actual systems can be explained by chance; otherwise, the difference is significant at level α. Following MUC's example, we set α to 0.1. A total of 11 teams participated in the smoking challenge. The training data for the challenge were released in July 2006, and the test data were released for only three days in September 2006. Each team was permitted to submit up to three system on the test A total of were count only one of the three of and their other two were evaluated from the (see and are not in this In this we describe each et a this classifier that to the smoking status of patients and then and to the information. The of the that only a few in a record to the smoking status and that these could be easily by their (e.g., If more than one in a record smoking then only the was If such were then the record was classified as To predict the smoking status of a record from the test set, each from this record was compared with from the training The of the measures between each and the most similar in the training set the smoking status of the et various for text with and found that the performance from the of and that and classification more than et also from lack of explicit smoking information in the Unknown and these to further two to processing the In the they classified based on to smoking using In the they classified each explicit to smoking using and then to categories from judgments of to smoking the smoking category for the collection the category in the Given the data of their et the i2b2 data set with records and their smoking categories. that a based on the data set perform better than trained on each of the data sets This hypothesis is with a the of a system is as the sample et smoking status evaluation system was with and from the Medical This system document (e.g., medical and the status of these entities (e.g., smoking-related such as and It thus the set of and of explicit to smoking. et found that the using on the to smoking better than the classification using (see and in Table and Table this system are by the of the and for example, to the system by et The and of explicit to their For example, the does not in the section could both to a Past Smoker and to a Non-Smoker, on However, such and information was to and for Precision, and by and for Precision, and by on and by not in macroaveraged not in microaveraged the is on and by not in macroaveraged not in microaveraged the is smoking status evaluation as a task and a process. The marked smoking-related In Cohen's a smoking-related was a of specific (e.g., and The the with the from the The without any specific and them The the and classified them using created of system by a system of for the smoking status of This system utilized both and such as of and of three one and two for the smoking status of patients of smoking challenge systems as a to the and can be at For with from the and found decision the most evaluated using on the training this classifier its performance trained using only the and their as For with and were with up to five words, one of the two in the with the These two systems to most of the test data to the Unknown category. found that of as well as Table shows that from in both microaveraged and macroaveraged F-measures, at α = 0.1. the Medical to the challenge. is a system for medical text processing in and records by For the smoking challenge, this system was to English medical discharge et the smoking challenge as a classification system marked the smoking status category of each in a record and to classify the document The judgments of the training set were derived Most of the text processing was through the of the Information system of and classification was in three and by using with an In the the Unknown category was a of that excluded this category. In the the category was through the of For with was In the the Smoker, and Past Smoker categories were by a of to Current and Past This was as a temporal and was using an the in this the were not so as to information of such as the of the and such as and between Current and Past the categories were The document categories were from to as follows: Current Smoker, Past Smoker, Smoker, Non-Smoker, and Each document was the of the category for which it could provide For example, a Current Smoker was to a document of any as Current et and the system to the smoking challenge. is designed for and clinical information from free text, and contains a patient smoking status This system of four document and The and the of a record based on the of the section The to them with in a The the text and is to The to labels to each The set of is then using provides based on of patient This system for a given patient and a to the with the to each To the differences of discharge summaries from patient et to the smoking challenge. this they into the differences in the labels of the i2b2 smoking challenge and the smoking they the system to a specific than a set of and for each The reported that they found the temporal of smoking categories to be a significant challenge. et of the explicit to smoking status. and using the results of these and records using a In addition to they used and information about and used a system to predict the smoking status based on explicit of smoking. that the system they for information about the smoking status of patients the records were of explicit smoking-related For they all explicit of smoking from the records and created a data set that only the trained two on the showed that annotators precision, recall, and F-measure on the The two systems the performance of annotators with of on the test Table and 1 show the precision, recall, and F-measure, both macroaveraged and for each of the system to the smoking challenge. Table shows the results of the significance on microaveraged and macroaveraged of the this reveals that the differences in the microaveraged of the systems are not significant at α = of these systems are not from each other in their macroaveraged at this α. Results from Table by microaveraged Table shows that systems that used similar machine e.g., the systems that used not all perform This that other such as the and the used with the also to the Table also shows that a majority of the performances from systems that of the lack of explicit smoking information in the Unknown records and to further processing or classification (see the results of et and et These systems microaveraged to that were on explicit to smoking status and et of the systems that the smoking classification was Unknown unless some in the document showed it to be this in an F-measure of in the Unknown category for and et and et Table 6 shows the precision, recall, and F-measure of each system on each of the five smoking status categories. In the systems successfully the Unknown and categories; some Current and Past with the of the systems of et systems correctly classified one of the three all in Precision, and for in Precision, and for in In addition to system performance on the test set and on categories in the test set, we the performance and agreement of systems on data i.e., records, in the test For we used κ. Table shows the level of κ agreement of each system with the ground truth as well as the level of κ agreement of pairs of Most is the level of agreement between systems by the same but only for some of the This level of agreement is not these from of the same The systems of et in their to and and only in their training and used a that the from and Agreement of and the Agreement between 0.8 and is in and agreement is in Agreement of and the Agreement between 0.8 and is in and agreement is in Table also shows in the systems showed a level of agreement with each For example, each of the systems of et with each of Cohen's systems showed κ between and systems showed strong agreement with each other the differences in their to smoking status For example, and system showed κ of to compared with et Similarly, the systems of et a κ of 0.8 with that of and the classified the records at the level whereas the was at the The level of agreement between to the smoking challenge, by the of these on the ground truth, indicates the of the set of to smoking status On the other the disagreements the systems their relative For example, and disagreed with each other on records = These systems showed relative strengths in in and Current Smoker, and in Unknown and Past were only three records that system marked The state of the art in smoking status evaluation be by the strengths of such of the teams with the annotation of a few of the in the challenge mentioned in the Annotations our medical records were annotated by pulmonologists. they provide a ground truth, judgments can However, given the medical of the annotators and the agreement on the labels of these records the we the annotations as the English of the records is for all of the a majority of the systems found the classification disagreed with the judgments of we that of should be by an to measure the of the that the in the of the ground truth only on our findings on the smoking challenge, we to our in two we are by the strengths of the systems and to of these to improve the we and with data to be We that a system provide insights into the intuitive judgments on the smoking status of In this we the i2b2 smoking challenge, the data and the data preparation the evaluation and each of the We the evaluation and of the system and provided a of the of our findings for For this challenge, we a document collection derived from actual medical discharge records, this collection in to a real-world medical classification and system performances. We showed that asked to a decision on the smoking status of patients based on the explicitly stated information in medical discharge annotators with each other more than of the The systems that participated in the smoking challenge represented various from and to the differences in their to smoking status many of these systems In there were system with microaveraged above A majority of these systems of the of the challenge data, e.g., lack of to smoking in records marked of explicit to smoking. the systems in the smoking challenge showed that discharge summaries smoking status using a limited number of textual (e.g., of the smoking status from these The authors all teams for their to the challenge, for their in the of the workshop that the challenge, and and the of for their on this Özlem Uzuner, Ira Goldstein, Yuan Luo 0001, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 1 |
| 2007 | Three Approaches to Automatic Assignment of ICD-9-CM Codes to Radiology Reports
Ira Goldstein, Anna Arzumtsyan, Özlem Uzuner |
AMIA | 3 |
| 2007 | Viewpoint Paper: Evaluating the State-of-the-Art in Automatic De-identificationabstractTo facilitate and survey studies in automatic de-identification, as a part of the i2b2 (Informatics for Integrating Biology to the Bedside) project, authors organized a Natural Language Processing (NLP) challenge on automatically removing private health information (PHI) from medical discharge records. This manuscript provides an overview of this de-identification challenge, describes the data and the annotation process, explains the evaluation metrics, discusses the nature of the systems that addressed the challenge, analyzes the results of received system runs, and identifies directions for future research. The de-indentification challenge data consisted of discharge summaries drawn from the Partners Healthcare system. Authors prepared this data for the challenge by replacing authentic PHI with synthesized surrogates. To focus the challenge on non-dictionary-based de-identification methods, the data was enriched with out-of-vocabulary PHI surrogates, i.e., made up names. The data also included some PHI surrogates that were ambiguous with medical non-PHI terms. A total of seven teams participated in the challenge. Each team submitted up to three system runs, for a total of sixteen submissions. The authors used precision, recall, and F-measure to evaluate the submitted system runs based on their token-level and instance-level performance on the ground truth. The systems with the best performance scored above 98% in F-measure for all categories of PHI. Most out-of-vocabulary PHI could be identified accurately. However, identifying ambiguous PHI proved challenging. The performance of systems on the test data set is encouraging. Future evaluations of these systems will involve larger data sets from more heterogeneous sources. Özlem Uzuner, Yuan Luo 0001, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 1 |
| 2006 | Syntactically-Informed Semantic Category Recognizer for Discharge Summaries
Tawanda C. Sibanda, Peter Szolovits, Özlem Uzuner |
AMIA | 4 |
| 2006 | Role of Local Context in Automatic Deidentification of Ungrammatical, Fragmented Text
Tawanda C. Sibanda, Özlem Uzuner |
HLT-NAACL | 2 |
| 2005 | Capturing Expression Using Linguistic Information
Özlem Uzuner, Boris Katz |
AAAI | 1 |
| 2005 | A Comparative Study of Language Models for Book and Author Recognition
Özlem Uzuner, Boris Katz |
IJCNLP | 1 |
| 2003 | Content and expression-based copy recognition for intellectual property protectionabstractProtection of copyrights and revenues of content owners in the digital world has been gaining importance in the recent years. This paper presents a way of fingerprinting text documents that can be used to identify content and expression similarities in documents, as a way of facilitating tracking of digital copies of works, to ensure proper compensation to content owners. Özlem Uzuner, Randall Davis |
Digital Rights Management Workshop | 1 |