VLDB 2026 Research / reviewers in the wild / expert
Wendy W. Chapman
dblp:70/2019 · also Wendy Webber Chapman
· DBLP profile ↗
96ranked-venue papers
18as first author
5since 2021 · last 2023
0000-0001-8702-4483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 94 · 17 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Developing a deep learning natural language processing algorithm for automated reporting of adverse drug reactions
Christopher McMaster, Julia Chan, David F. L. Liew, Elizabeth Su, Albert G. Frauman, Wendy W. Chapman, Douglas E. V. Pires |
J. Biomed. Informatics | 6 |
| 2022 | The Validitron Sandbox: a cloud environment to support prototyping, workflow design and integration testing of data-focused digital health applications
Kit Huckvale, Wendy W. Chapman, Daniel Capurro |
AMIA | 2 |
| 2021 | Impact of Comorbidity Profiles on Pain Trajectories in Breast Cancer Patients by Using Electronic Health Record Data
Jia-Wen Guo, Katherine A. Sward, Ann M. Lyons, Susan L. Beck, Gary W. Donaldson, Wendy W. Chapman, Lewis J. Frey |
AMIA | 6 |
| 2021 | Design and evaluation of a Women in American Medical Informatics Association (AMIA) leadership programabstractThe objective is to report on the design and evaluation of the inaugural Women in AMIA Leadership Program. A year-long leadership curriculum was developed. Survey responses were summarized with descriptive statistics and quotes selected. Twenty-four scholars participated in the program. There was a significant increase in perceived achievement of learning objectives after the program (P < .0001). The largest improvement was in leadership confidence and presence in work interactions (modal answer Neutral in presurvey from 21 responses rose to Agree in postsurvey from 24 responses). Most (92% of 13) scholars clarified leadership vision and goals and (83% of 18) would be Very Likely to recommend the program to others. The goals of the program-developing women's leader identity, increasing networks, and accumulating experience for future programs-were achieved. The second leadership program is on its way in the United States and Australia. This study may benefit organizations seeking to develop leadership programs for women in informatics and digital health. María Adela Grando, Jessica S. Ancker, Donghua Tao, Rachael Howe, Clare Coonan, Merida L. Johns, Wendy W. Chapman |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Adaptation of an NLP system to a new healthcare environment to identify social determinants of health
Ruth M. Reeves, Lee M. Christensen, Jeremiah R. Brown, Michael Conway, Maxwell Levis, Glenn T. Gobbel, Rashmee U. Shah, Christine Goodrich, Iben Ricket, Freneka F. Minter, Andrew Bohm, Bruce E. Bray, Michael E. Matheny, Wendy W. Chapman |
J. Biomed. Informatics | 14 |
| 2020 | Training Women for Leadership in Informatics and Digital Health: A Report from the Inaugural Women in AMIA Leadership Program
Wendy W. Chapman, María Adela Grando, Merida L. Johns, Guergana K. Savova, Maia Hightower |
AMIA | 1 |
| 2020 | A Corpus Analysis of Social Isolation from Clinical Notes of Patients with Cancer
Jia-Wen Guo, Christina L. Radloff, Katherine A. Sward, Susan L. Beck, Wendy W. Chapman, Gary W. Donaldson, Lewis J. Frey |
AMIA | 5 |
| 2019 | Determination of Marital Status of Patients from Structured and Unstructured Electronic Healthcare Data
Brian T. Bucher, Jianlin Shi, Robert J. Pettit, Jeffrey P. Ferraro, Wendy W. Chapman, Adi V. Gundlapalli |
AMIA | 5 |
| 2019 | Researchers' Perspectives on Symptoms related to Cancer Pain: A Network Analysis of Literatures
Jia-Wen Guo, Christina L. Radloff, Susan L. Beck, Gary W. Donaldson, Wendy W. Chapman, Katherine A. Sward, Lewis J. Frey |
AMIA | 5 |
| 2019 | Extracting Disease Onset from Family History Comments in the Electronic Health Record using Fast Healthcare Interoperability Resources
Jianlin Shi, Kensaku Kawamoto, Wendy Kohlmann, Danielle L. Mowery, Richard L. Bradshaw, Subhadeep Deep, Wendy W. Chapman, Guilherme Del Fiol |
AMIA | 7 |
| 2019 | Using Natural Language Processing to improve EHR Structured Data-based Surgical Site Infection Surveillance
Jianlin Shi, Siru Liu, Liese C. Pruitt, Carolyn Luppens, Jeffrey P. Ferraro, Adi V. Gundlapalli, Wendy W. Chapman, Brian T. Bucher |
AMIA | 7 |
| 2018 | Natural Language Processing at Scale - Perspectives from Five Healthcare Organizations
Rebecca S. Jacobson, Wendy W. Chapman, Olga V. Patterson, Daniel Zisook |
AMIA | 2 |
| 2018 | Detection of Healthcare-Associated Infections Using Electronic Health Record Data
Siru Liu, Jeffrey P. Ferraro, Adi V. Gundlapalli, Wendy W. Chapman, Brian T. Bucher |
AMIA | 4 |
| 2018 | Detecting Current Episodes of Cholecystitis-related Pain from Veterans Affairs Clinical Notes using Natural Language Processing
Danielle L. Mowery, Luke Martin, Brett R. South, Ellen Morrow, William Peche, Eric Wiesner, Wendy W. Chapman, Benjamin S. Brooke |
AMIA | 7 |
| 2018 | Annotating Social Determinants of Health and Functional Status Information Using Publicly Accessible Corpora
Ruth M. Reeves, Brett R. South, Glenn T. Gobbel, Lee M. Christensen, Wendy W. Chapman, Michael E. Matheny, Jeremiah R. Brown |
AMIA | 6 |
| 2018 | NLPReViz: an interactive tool for natural language processing on clinical textabstractThe gap between domain experts and natural language processing expertise is a barrier to extracting understanding from clinical text. We describe a prototype tool for interactive review and revision of natural language processing models of binary concepts extracted from clinical notes. We evaluated our prototype in a user study involving 9 physicians, who used our tool to build and revise models for 2 colonoscopy quality variables. We report changes in performance relative to the quantity of feedback. Using initial training sets as small as 10 documents, expert review led to final F1scores for the "appendiceal-orifice" variable between 0.78 and 0.91 (with improvements ranging from 13.26% to 29.90%). F1for "biopsy" ranged between 0.88 and 0.94 (-1.52% to 11.74% improvements). The average System Usability Scale score was 70.56. Subjective feedback also suggests possible design improvements. Gaurav Trivedi, Phuong Pham, Wendy W. Chapman, Rebecca Hwa, Janyce Wiebe, Harry Hochheiser |
J. Am. Medical Informatics Assoc. | 3 |
| 2018 | Using clinical Natural Language Processing for health outcomes research: Overview and actionable suggestions for future advancesabstractThe importance of incorporating Natural Language Processing (NLP) methods in clinical informatics research has been increasingly recognized over the past years, and has led to transformative advances. Typically, clinical NLP systems are developed and evaluated on word, sentence, or document level annotations that model specific attributes and features, such as document content (e.g., patient status, or report type), document section types (e.g., current medications, past medical history, or discharge summary), named entities and concepts (e.g., diagnoses, symptoms, or treatments) or semantic attributes (e.g., negation, severity, or temporality). From a clinical perspective, on the other hand, research studies are typically modelled and evaluated on a patient- or population-level, such as predicting how a patient group might respond to specific treatments or patient monitoring over time. While some NLP tasks consider predictions at the individual or group user level, these tasks still constitute a minority. Owing to the discrepancy between scientific objectives of each field, and because of differences in methodological evaluation priorities, there is no clear alignment between these evaluation approaches. Here we provide a broad summary and outline of the challenging issues involved in defining appropriate intrinsic and extrinsic evaluation methods for NLP research that is to be used for clinical outcomes research, and vice versa. A particular focus is placed on mental health research, an area still relatively understudied by the clinical NLP research community, but where NLP methods are of notable relevance. Recent advances in clinical NLP method development have been significant, but we propose more emphasis needs to be placed on rigorous evaluation for the field to advance further. To enable this, we provide actionable suggestions, including a minimal protocol that could be used when reporting clinical NLP method development and its evaluation. Sumithra Velupillai, Hanna Suominen, Maria Liakata, Angus Roberts, Anoop D. Shah, Katherine Morley, David Osborn, Joseph Hayes, Robert Stewart 0002, Johnny Downs, Wendy W. Chapman, Rina Dutta |
J. Biomed. Informatics | 11 |
| 2017 | "My work will surely speak for itself: " Visibility, Networking, and Self Promotion in Informatics
Wendy W. Chapman, Murielle S. Beene, Omolola Ogunyemi, Genevieve B. Melton, Laura K. Wiley |
AMIA | 1 |
| 2017 | Detecting Evidence of Intra-abdominal Surgical Site Infections from Radiology Reports Using Natural Language Processing
Alec B. Chapman, Danielle L. Mowery, Douglas S. Swords, Wendy W. Chapman, Brian T. Bucher |
AMIA | 4 |
| 2017 | A Comparison of Stroke Classifiers Leveraging Hospital Billing Codes versus Natural Language Processing
Danielle L. Mowery, Brent D. Hill, Wendy W. Chapman, Lisa A. Cannon-Albright, Jennifer Majersik |
AMIA | 4 |
| 2016 | Understanding patient satisfaction with received healthcare services: A natural language processing approach
Kristina Doing-Harris, Danielle L. Mowery, Chrissy Daniels, Wendy W. Chapman, Mike Conway |
AMIA | 4 |
| 2016 | VIP: A Framework for Mining Clinical Concepts Using Knowledge Author Ontologies and Apache UIMA
Thomas Ginter, Lalindra De Silva, Olga V. Patterson, William Scuba, Wendy W. Chapman, Scott L. DuVall |
AMIA | 5 |
| 2016 | Women in Informatics Leadership Forum
Rebecca S. Jacobson, Suzanne Bakken, Wendy W. Chapman, Valerie Florance, Jessica D. Tenenbaum |
AMIA | 3 |
| 2016 | A Characterization of Emotional Valence to Support Review of Free-Text Press-Ganey Patient Satisfaction Survey Responses
Danielle L. Mowery, Kristina Doing-Harris, Wendy W. Chapman, Chrissy Daniels, Mike Conway |
AMIA | 3 |
| 2016 | Event Coreference in Support of Temporal Reasoning in Mental Health Notes
Ruth M. Reeves, Marcus Verhagen, Cynthia Brandt, Wendy W. Chapman, Michael E. Matheny, Steven H. Brown, Brian Marx, Theodore Speroff |
AMIA | 4 |
| 2015 | RapTAT: A Tool for Assisted Annotation and Reviewer Training via Online Machine Learning
Glenn T. Gobbel, Ruth M. Reeves, Brett R. South, Wendy W. Chapman, Jay Jarman, Steven Lay, Michael E. Matheny |
AMIA | 4 |
| 2015 | Developing Natural Language Processing Systems for Healthcare
Ruth M. Reeves, Wendy W. Chapman, Dezon Finch, Jennifer H. Garvin, Glenn T. Gobbel |
AMIA | 2 |
| 2015 | Annotating ADLs and IADLs in Veterans Affairs Clinical Documents
Brett R. South, Danielle L. Mowery, Lee M. Christensen, Adi V. Gundlapalli, Melissa Tharp, Marzieh Vali, Marjorie Carter, Mike Conway, Salomeh Keyhani, Wendy W. Chapman |
AMIA | 10 |
| 2015 | Towards a Generalizable Time Expression Model for Temporal Reasoning in Clinical Notes
Sumithra Velupillai, Danielle L. Mowery, Samir E. AbdelRahman, Lee M. Christensen, Wendy W. Chapman |
AMIA | 5 |
| 2015 | Evaluating the state of the art in disorder recognition and normalization of the clinical narrativeabstractOBJECTIVE: The ShARe/CLEF eHealth 2013 Evaluation Lab Task 1 was organized to evaluate the state of the art on the clinical text in (i) disorder mention identification/recognition based on Unified Medical Language System (UMLS) definition (Task 1a) and (ii) disorder mention normalization to an ontology (Task 1b). Such a community evaluation has not been previously executed. Task 1a included a total of 22 system submissions, and Task 1b included 17. Most of the systems employed a combination of rules and machine learners. MATERIALS AND METHODS: We used a subset of the Shared Annotated Resources (ShARe) corpus of annotated clinical text--199 clinical notes for training and 99 for testing (roughly 180 K words in total). We provided the community with the annotated gold standard training documents to build systems to identify and normalize disorder mentions. The systems were tested on a held-out gold standard test set to measure their performance. RESULTS: For Task 1a, the best-performing system achieved an F1 score of 0.75 (0.80 precision; 0.71 recall). For Task 1b, another system performed best with an accuracy of 0.59. DISCUSSION: Most of the participating systems used a hybrid approach by supplementing machine-learning algorithms with features generated by rules and gazetteers created from the training data and from external resources. CONCLUSIONS: The task of disorder normalization is more challenging than that of identification. The ShARe corpus is available to the community as a reference standard for future studies. Sameer Pradhan, Noémie Elhadad, Brett R. South, David Martínez 0001, Lee M. Christensen, Amy Vogel, Hanna Suominen, Wendy W. Chapman, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 8 |
| 2014 | Developing a Knowledge Base for Detecting Carotid Stenosis with pyConText
Danielle L. Mowery, Daniel Franc, Shazia Ashfaq, Eric Cheng, Tania Zamora, Wendy W. Chapman, Brian E. Chapman |
AMIA | 6 |
| 2014 | A System Usability Study Assessing a Machine-Assisted Interactive Interface to Support Annotation of Protected Health Information in Clinical Texts
Brett R. South, Danielle L. Mowery, Chris Leng, Stéphane M. Meystre, Wendy W. Chapman |
AMIA | 5 |
| 2014 | Disease/Disorder Semantic Template Filling - Information Extraction Challenge in the ShARe/CLEF eHealth Evaluation Lab 2014
Sumithra Velupillai, Danielle L. Mowery, Lee M. Christensen, Noémie Elhadad, Sameer Pradhan, Guergana K. Savova, Wendy W. Chapman |
AMIA | 7 |
| 2014 | Cue-based assertion classification for Swedish clinical text - Developing a lexicon for pyConTextSweabstractOBJECTIVE: The ability of a cue-based system to accurately assert whether a disorder is affirmed, negated, or uncertain is dependent, in part, on its cue lexicon. In this paper, we continue our study of porting an assertion system (pyConTextNLP) from English to Swedish (pyConTextSwe) by creating an optimized assertion lexicon for clinical Swedish. METHODS AND MATERIAL: We integrated cues from four external lexicons, along with generated inflections and combinations. We used subsets of a clinical corpus in Swedish. We applied four assertion classes (definite existence, probable existence, probable negated existence and definite negated existence) and two binary classes (existence yes/no and uncertainty yes/no) to pyConTextSwe. We compared pyConTextSwe's performance with and without the added cues on a development set, and improved the lexicon further after an error analysis. On a separate evaluation set, we calculated the system's final performance. RESULTS: Following integration steps, we added 454 cues to pyConTextSwe. The optimized lexicon developed after an error analysis resulted in statistically significant improvements on the development set (83% F-score, overall). The system's final F-scores on an evaluation set were 81% (overall). For the individual assertion classes, F-score results were 88% (definite existence), 81% (probable existence), 55% (probable negated existence), and 63% (definite negated existence). For the binary classifications existence yes/no and uncertainty yes/no, final system performance was 97%/87% and 78%/86% F-score, respectively. CONCLUSIONS: We have successfully ported pyConTextNLP to Swedish (pyConTextSwe). We have created an extensive and useful assertion lexicon for Swedish clinical text, which could form a valuable resource for similar studies, and which is publicly available. Sumithra Velupillai, Maria Skeppstedt, Maria Kvist, Danielle L. Mowery, Brian E. Chapman, Hercules Dalianis, Wendy W. Chapman |
Artif. Intell. Medicine | 7 |
| 2014 | Evaluating the effects of machine pre-annotation and an interactive annotation interface on manual de-identification of clinical textabstractThe Health Insurance Portability and Accountability Act (HIPAA) Safe Harbor method requires removal of 18 types of protected health information (PHI) from clinical documents to be considered "de-identified" prior to use for research purposes. Human review of PHI elements from a large corpus of clinical documents can be tedious and error-prone. Indeed, multiple annotators may be required to consistently redact information that represents each PHI class. Automated de-identification has the potential to improve annotation quality and reduce annotation time. For instance, using machine-assisted annotation by combining de-identification system outputs used as pre-annotations and an interactive annotation interface to provide annotators with PHI annotations for "curation" rather than manual annotation from "scratch" on raw clinical documents. In order to assess whether machine-assisted annotation improves the reliability and accuracy of the reference standard quality and reduces annotation effort, we conducted an annotation experiment. In this annotation study, we assessed the generalizability of the VA Consortium for Healthcare Informatics Research (CHIR) annotation schema and guidelines applied to a corpus of publicly available clinical documents called MTSamples. Specifically, our goals were to (1) characterize a heterogeneous corpus of clinical documents manually annotated for risk-ranked PHI and other annotation types (clinical eponyms and person relations), (2) evaluate how well annotators apply the CHIR schema to the heterogeneous corpus, (3) compare whether machine-assisted annotation (experiment) improves annotation quality and reduces annotation time compared to manual annotation (control), and (4) assess the change in quality of reference standard coverage with each added annotator's annotations. Brett R. South, Danielle L. Mowery, Ying Suo, Jianwei Leng, Óscar Ferrández, Stéphane M. Meystre, Wendy W. Chapman |
J. Biomed. Informatics | 7 |
| 2013 | Identifying Synonymy between SNOMED Clinical Terms of Varying Length Using Distributional Analysis of Electronic Health Records
Aron Henriksson, Mike Conway, Martin Duneld, Wendy W. Chapman |
AMIA | 4 |
| 2013 | An Information Extraction Based Search System for Clinical Records
Yan Jiao, Wendy W. Chapman |
AMIA | 3 |
| 2013 | Schema Builder: A Web-based User Interface for Authoring and Sharing Natural-Language Processing Schemas
Melissa Tharp, Matthew K. Hong, Harry Hochheiser, Wendy W. Chapman |
AMIA | 5 |
| 2013 | Semantic Annotation of Clinical Events for Generating a Problem List
Danielle L. Mowery, Pamela W. Jordan, Janyce Wiebe, Henk Harkema, Wendy W. Chapman |
AMIA | 5 |
| 2013 | Creating a Reference Standard of Acronym and Abbreviation Annotations for the ShARe/CLEF eHealth Challenge 2013
Danielle L. Mowery, Brett R. South, Jianwei Leng, Laura-Maria Peltonen, Riitta Danielsson-Ojala, Sanna Salanterä, Wendy W. Chapman |
AMIA | 7 |
| 2013 | Panel: Shared Resources, Shared Code, and Shared Activities in Clinical Natural Language Processing
Guergana K. Savova, Wendy W. Chapman, Noémie Elhadad, Martha Palmer |
AMIA | 2 |
| 2013 | Knowledge Representation for Topic Model Based Discharge Summary Clustering
Yan Jiao, Wendy W. Chapman |
AMIA | 3 |
| 2013 | Improving performance of natural language processing part-of-speech tagging on clinical narratives through domain adaptationabstractOBJECTIVE: Natural language processing (NLP) tasks are commonly decomposed into subtasks, chained together to form processing pipelines. The residual error produced in these subtasks propagates, adversely affecting the end objectives. Limited availability of annotated clinical data remains a barrier to reaching state-of-the-art operating characteristics using statistically based NLP tools in the clinical domain. Here we explore the unique linguistic constructions of clinical texts and demonstrate the loss in operating characteristics when out-of-the-box part-of-speech (POS) tagging tools are applied to the clinical domain. We test a domain adaptation approach integrating a novel lexical-generation probability rule used in a transformation-based learner to boost POS performance on clinical narratives. METHODS: Two target corpora from independent healthcare institutions were constructed from high frequency clinical narratives. Four leading POS taggers with their out-of-the-box models trained from general English and biomedical abstracts were evaluated against these clinical corpora. A high performing domain adaptation method, Easy Adapt, was compared to our newly proposed method ClinAdapt. RESULTS: The evaluated POS taggers drop in accuracy by 8.5-15% when tested on clinical narratives. The highest performing tagger reports an accuracy of 88.6%. Domain adaptation with Easy Adapt reports accuracies of 88.3-91.0% on clinical texts. ClinAdapt reports 93.2-93.9%. CONCLUSIONS: ClinAdapt successfully boosts POS tagging performance through domain adaptation requiring a modest amount of annotated clinical data. Improving the performance of critical NLP subtasks is expected to reduce pipeline error propagation leading to better overall results on complex processing tasks. Jeffrey P. Ferraro, Hal Daumé III, Scott L. DuVall, Wendy W. Chapman, Henk Harkema, Peter J. Haug |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Using chief complaints for syndromic surveillance: A review of chief complaint based classifiers in North America
Mike Conway, John N. Dowling, Wendy W. Chapman |
J. Biomed. Informatics | 3 |
| 2012 | A Preliminary Approach for Creating a Semi-synthetic Multimodal Clinical Data Set from a Publicly Available Image Repository
Shazia Ashfaq, Amilcare Gentili, Wendy W. Chapman, Brian E. Chapman |
AMIA | 3 |
| 2012 | Discovering Lexical Instantiations of Clinical Concepts using Web Services, WordNet and Corpus Resources
Mike Conway, Wendy W. Chapman |
AMIA | 2 |
| 2012 | TxtVect: A Tool for Extracting Features from Clinical Documents
Wendy W. Chapman |
AMIA | 2 |
| 2012 | On the Road Towards Developing a Publicly Available Corpus of De-identified Clinical Texts
Brett R. South, Danielle L. Mowery, Óscar Ferrández, Shuying Shen, Ying Suo, Annie Chen, Stéphane M. Meystre, Wendy W. Chapman |
AMIA | 10 |
| 2012 | iDASH: integrating data for analysis, anonymization, and sharingabstractiDASH (integrating data for analysis, anonymization, and sharing) is the newest National Center for Biomedical Computing funded by the NIH. It focuses on algorithms and tools for sharing data in a privacy-preserving manner. Foundational privacy technology research performed within iDASH is coupled with innovative engineering for collaborative tool development and data-sharing capabilities in a private Health Insurance Portability and Accountability Act (HIPAA)-certified cloud. Driving Biological Projects, which span different biological levels (from molecules to individuals to populations) and focus on various health conditions, help guide research and development within this Center. Furthermore, training and dissemination efforts connect the Center with its stakeholders and educate data owners and data consumers on how to share and use clinical and biological data. Through these various mechanisms, iDASH implements its goal of providing biomedical and behavioral researchers with access to data, software, and a high-performance computing environment, thus enabling them to generate and test new hypotheses. Lucila Ohno-Machado, Vineet Bafna, Aziz A. Boxwala, Brian E. Chapman, Wendy W. Chapman, Kamalika Chaudhuri, Michele E. Day, Claudiu Farcas, Nathaniel D. Heintzman, Xiaoqian Jiang, Hyeon-Eui Kim, Jihoon Kim 0001, Michael E. Matheny, Frederic S. Resnic, Staal Amund Vinterbo |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | A system for coreference resolution for the clinical narrativeabstractOBJECTIVE: To research computational methods for coreference resolution in the clinical narrative and build a system implementing the best methods. METHODS: The Ontology Development and Information Extraction corpus annotated for coreference relations consists of 7214 coreferential markables, forming 5992 pairs and 1304 chains. We trained classifiers with semantic, syntactic, and surface features pruned by feature selection. For the three system components--for the resolution of relative pronouns, personal pronouns, and noun phrases--we experimented with support vector machines with linear and radial basis function (RBF) kernels, decision trees, and perceptrons. Evaluation of algorithms and varied feature sets was performed using standard metrics. RESULTS: The best performing combination is support vector machines with an RBF kernel and all features (MUC score=0.352, B(3)=0.690, CEAF=0.486, BLANC=0.596) outperforming a traditional decision tree baseline. DISCUSSION: The application showed good performance similar to performance on general English text. The main error source was sentence distances exceeding a window of 10 sentences between markables. A possible solution to this problem is hinted at by the fact that coreferent markables sometimes occurred in predictable (although distant) note sections. Another system limitation is failure to fully utilize synonymy and ontological knowledge. Future work will investigate additional ways to incorporate syntactic features into the coreference problem. CONCLUSION: We investigated computational methods for coreference resolution in the clinical narrative. The best methods are released as modules of the open source Clinical Text Analysis and Knowledge Extraction System and Ontology Development and Information Extraction platforms. Jiaping Zheng, Wendy W. Chapman, Timothy A. Miller, Chen Lin 0002, Rebecca S. Jacobson, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Anaphoric reference in clinical reports: Characteristics of an annotated corpus
Wendy W. Chapman, Guergana K. Savova, Jiaping Zheng, Melissa Tharp, Rebecca S. Jacobson |
J. Biomed. Informatics | 1 |
| 2012 | An approach to improve LOINC mapping through augmentation of local test names
Hyeon-Eui Kim, Robert El-Kareh, Anupam Goel, F. N. U. Vineet, Wendy W. Chapman |
J. Biomed. Informatics | 5 |
| 2012 | Building an automated SOAP classifier for emergency department reports
Danielle L. Mowery, Janyce Wiebe, Shyam Visweswaran, Henk Harkema, Wendy W. Chapman |
J. Biomed. Informatics | 5 |
| 2011 | Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutionsabstractThis issue of JAMIA focuses on natural language processing (NLP) techniques for clinical-text information extraction. Several articles are offshoots of the yearly ‘Informatics for Integrating Biology and the Bedside’ (i2b2) (http://www.i2b2.org) NLP shared-task challenge, introduced by Uzuner et al (see page 552)1 and co-sponsored by the Veteran's Administration for the last 2 years. This shared task follows long-running challenge evaluations in other fields, such as the Message Understanding Conference (MUC) for information extraction,2 TREC3 for text information retrieval, and CASP4 for protein structure prediction. Shared tasks in the clinical domain are recent and include annual i2b2 Challenges that began in 2006, a challenge for multi-label classification of radiology reports sponsored by Cincinnati Children's Hospital in 2007,5 a 2011 Cincinnati Children's Hospital challenge on suicide notes,6 and the 2011 TREC information retrieval shared task involving retrieval of clinical cases from narrative records.7 Although NLP research in the clinical domain has been active since the 1960s, progress in the development of NLP applications for clinical text has been slow and lags behind progress in the general NLP domain. There are several barriers to NLP development in the clinical domain, and shared tasks like the i2b2/VA Challenge address some of these barriers. Nevertheless, many barriers remain and unless the community takes a more active role in developing novel approaches for addressing the barriers, advancement and innovation will continue to be slow. Historically, there have been substantial barriers to NLP development in the clinical domain. These barriers are not unique to the clinical domain: they also occur in the fields of software engineering and general NLP. Because of concerns regarding patient privacy and worry about revealing unfavorable institutional practices, hospitals and clinics have been extremely reluctant to allow access to clinical data for researchers from outside the associated institutions. The lack of reliable and inexpensive de-identification techniques for narrative reports has compounded the reluctance to share. Such restricted access to shared datasets has hindered collaboration and inhibited the ability to assess and adapt NLP technologies across institutions and among research groups. Several pioneering efforts5,8–11 have made clinical data available for sharing—we need more of these grass-roots efforts. Closely related but not completely conditional on lack of shared datasets is the deficiency of annotated clinical data for training NLP applications and benchmarking performance. The sublanguage of clinical reports often necessitates domain-specific development and training, and, as a consequence, NLP modules developed for general text typically do not perform as well on clinical narratives. We need increased coordination to create annotation sets that can be merged to produce larger training and evaluation sets. Without the ability to share data, the community has lacked incentives for developing common data models for manual and automatic annotations. The result is that annotated datasets are usually unique to the laboratory that generated them and thus remain small and that NLP modules that perform the same tasks cannot be substituted and compared without considerable translational effort. At present, the clinical NLP community is leveraging existing standards and conventions and working together to develop shared data models and to map annotations across information extraction applications. Adopting an existing NLP application or module is complicated—source code and documentation may be unavailable, and published descriptions may lack sufficient detail for reproducibility. Open source releases of clinical information extraction and retrieval systems have improved the opportunity to reproduce performance.12–15 Even with open source release, a tool may work less well in others' hands than in the hands of the original developers. Compounding the problem of reproducibility is the fact that proof-of-concept tools created in academic/research environments may not meet the highest software engineering quality, maintainability, scalability, or usability standards. And sometimes a tool may be over-fitted to a particular application, and modification to solve a similar problem may require wholesale changes. As Pedersen asserted,16 the NLP community needs to invest more in assisting others in applying and reproducing our results. In part due to previously listed barriers, collaboration within the clinical NLP community has been nominal. Development of NLP systems within the academic environment has centered around single institutions and single laboratories, and rather than building upon the foundations of previous work, the majority of clinical NLP systems developed over the last four decades have been reinvented as silos that are neither expanded nor applied outside of the individual laboratory. Other factors limiting collaboration include insufficient infrastructure for facilitating cooperation and the reality that collaboration is inherently inefficient. Nevertheless, as with the biomedical research community at large, a surge in progression beyond the last half century of research can only come through enhanced teamwork. A recent trend for teamwork across NLP research laboratories is evident in funded initiatives such as the VA's Consortium for Healthcare Informatics Research (CHIR)17 and in the ONC-funded SHARP Area 4 grant for Secondary Use of EHR data.18 Also, advances in open source development and conformity to common frameworks has led to recent advances in NLP allowing one team to extend work by another (eg, HiTEX12 built on GATE and cTAKES,14 ODIE,19 and Automated Retrieval Console (ARC) built on UIMA16). Although we are improving incrementally the predictive performance of clinical NLP tools, clinical NLP applications are seldom deployed in clinical, public health, or health services research settings. Currently, the perceived cost of applying NLP outweighs the perceived benefit. Deploying an NLP system typically requires a substantial amount of time from an expert NLP developer—normally, applications do not generalize and must be rebuilt, retrained, enhanced, and re-evaluated for each new task; the output of an NLP system typically requires extensive mapping to the specific problem being addressed; and the ability to aid a user in customizing the application is generally inadequate (see page 544).20 We need a shift of focus from accuracy in one task to generalizabiliy across many and from the production of papers as the sole output to production of usable software for medically relevant applications. We also need to understand where NLP tools fit into an overall user workflow so that the tools can be integrated into end-to-end applications for clinical, public health, and clinical research users. Shared tasks like the i2b2/VA Challenge address several of these barriers in part. Shared tasks provide annotated datasets to participants and sometimes to non-participants (i2b2 datasets are available to others a year after the Challenge). The i2b2 shared task is standardizing its corpus as much as possible—the same records are used from one year to the next with layers of annotation that build on each other, and common input/output specifications are applied every year. Shared tasks partially address the barrier of reproducibility by providing an evaluation opportunity that minimizes the risk of over-fitting: participants have time to train their systems in supervised fashion with an annotated training dataset, but then evaluation must be performed against a separate non-annotated dataset within a stringent time limit that prevents non-trivial system modifications. Although shared tasks are not designed for this purpose, the i2b2 Challenge has been the impetus for some new collaborations across independent research groups. Shared tasks have driven progress in related fields. For example, progress in speech understanding research was driven by a series of evaluations funded by DARPA from the late 1980s to the early 2000s.21 The research community was able to consistently drive down the error rate by a factor of two every 2 years, on successively more challenging tasks, moving from recognition of small-vocabulary read speech to automated transcription of broadcast news in multiple languages. Associated with this progress was the incorporation of speech recognition products into applications, from dictation to speech interfaces. Shared tasks provide value to the NLP community in several ways: Common evaluation metrics are developed. Annotated datasets are made available. Enticed by available annotated datasets, researchers in overlapping fields (both academic and corporate) participate in the tasks, bringing in new people and new approaches. Benchmarking evaluation on a shared dataset reveals the state-of-the-art performance for a given task. Students and post-docs receive excellent training opportunities. Preliminary results can be obtained by a new research group, which can potentially lead to funding opportunities. Pre-processed, standardized corpora with multiple layers of annotations on the same corpus pave the way for end-to-end evaluations in addition to evaluation on a single annotation layer. Conventions for standardizing annotations and input/output formats are developed, and despite other standardization efforts, shared task corpora often set de facto standards. In spite of the value of shared tasks, the tasks have several shortcomings: Participants come mainly from teams with funded projects that overlap with the shared task. For academic participants, a significant motivation is the opportunity to publish; however, there is sometimes limited value for the larger community in publications resulting from a shared task. Because development time is limited during shared tasks, participants often build on applications that already exist and apply methods already described in the literature. This can result in many similar approaches being applied to the same task. Although publishing the high-performing systems can be interesting, the resulting publications may not be novel and therefore may not improve the general body of knowledge. If a particular challenge task is repeated over time, there is a tendency for system approaches to converge on the approach that showed most success in the previous evaluation—evaluations repeated over time tend to reduce the diversity of approaches. Although shared tasks contribute to growth and progress, increased benefit to the community of clinical NLP developers and to potential users will require additional individual and community efforts that target existing barriers creatively. Driving progress in a way that will increase the impact of NLP in the realm of individual and population health will require creativity at both the grass-roots and the community levels. No single activity can tackle all barriers. In addition to encouraging variations on the development of shared tasks and their incentives, we would like to see new types of shared activities that foster the outcomes described below. In exchange for access to the costly annotated dataset, shared task participation could be contingent on depositing code in a shared repository or creating a web service for prospective users. In this model, the organizers could send test data to the participants' servers and the servers return the results for evaluation (see Leitner et al22 for a description of a metaserver used in evaluation of results from BioCreative II). The servers (and metaserver) could even persist, providing services to interested users beyond the initial shared task. Publication of computational methods in biomedical informatics journals like JAMIA could further encourage reproducibility of results through policies mandating simultaneous submission of code with a manuscript, as recommended by Pedersen.16 The clinical NLP community could independently accelerate reproducibility (and lead by example) by depositing code and developing web services in a common repository.23 This trend is occurring in settings restricted by affiliation, such as the VA VINCI framework24 (available only to VA researchers) and a cloud environment being hosted by the SHARP Area 4 grant18 (available to grant participants). The new National Center for Biomedical Computing iDASH25 is developing a similar cyber-infrastructure that will be restricted not by affiliation but by adherence to privacy policies and agreements required by data contributors. The National Library of Medicine is currently hosting a registry developed by the AMIA NLP working group called ORBIT for listing and pointing to biomedical informatics and NLP resources.26 In a shared task, dozens of research groups duplicate the same task independently. Although a variety of techniques can emerge for the same task, given the relatively short time frame allowed for development and training, the features and approaches applied in the challenge are often very similar. Whereas similarity of approaches reveals agreement among teams on the best approaches and sets the stage for collaboration, the competitive nature of shared tasks provides a disincentive to collaboration; the reward system for shared tasks is not at all dependent on the ability to collaborate across teams but is solely geared toward competition in which a single winner arises. The open source development community has found inherent rewards in collaborative development through supportive environments like GitHub. Perhaps we can learn from collaborative development communities who participate in hackathons27 and from the games industry in asking how a shared task can be designed so that collaboration is rewarded and becomes worthwhile, interesting, and attractive.28 Evaluation of a shared task is focused on accuracy, and existing challenges evaluate only predictive performance, not software engineering characteristics or usability. Imagine a shared task in which success is judged on usability of a system or direct portability of one technique to a new task or domain. Evaluating success of a system with this paradigm is inherently more complex, but we could learn from groupware evaluation, from the rich field of usability testing, and from incentivized competitions like those sponsored by the X-Prize Foundation.29 Because of the cost of creating annotated training data, shared tasks are often small scale, at least relative to real medical applications. We need new approaches to rapid adaptation of NLP systems to new applications, with less dependence on ‘deeply annotated’ data; such applications would present important opportunities for collaboration with the end user community, who might be motivated to provide domain expertise if they were likely to get a scalable, maintainable system out of the collaboration. Scalability will require more efficient techniques for manual annotation. And scalability will require an enriched ability to produce high quality software, which may necessitate better collaboration with industry30 and funding models that include support for operational development. The shared i2b2 evaluations have made a huge contribution to stimulating and vitalizing the field of clinical NLP; however, to ensure the transition into usable applications, the clinical NLP research community needs to address the critical issues of data access, development of shared infrastructure, and integration of software engineering methods to ensure the usability, maintainability, and availability of clinical NLP tools that are integrated into the workflow of real biomedical applications. This must be done in close collaboration with end users, software engineers, and clinical practitioners. We as a community need to think beyond the status quo of incremental improvement in the F score toward imaginative approaches that encourage collaboration, promote reproducibility, increase the scalability of NLP development, and provide value to end users. Authors are funded in part by U54HL108460, R01GM090187, SHARP ONC award 90TR0002, R01 CA127979, and U54LM008748. None. Not commissioned; internally peer reviewed. Wendy W. Chapman, Prakash M. Nadkarni, Lynette Hirschman, Leonard W. D'Avolio, Guergana K. Savova, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Developing a natural language processing application for measuring the quality of colonoscopy proceduresabstractOBJECTIVE: The quality of colonoscopy procedures for colorectal cancer screening is often inadequate and varies widely among physicians. Routine measurement of quality is limited by the costs of manual review of free-text patient charts. Our goal was to develop a natural language processing (NLP) application to measure colonoscopy quality. MATERIALS AND METHODS: Using a set of quality measures published by physician specialty societies, we implemented an NLP engine that extracts 21 variables for 19 quality measures from free-text colonoscopy and pathology reports. We evaluated the performance of the NLP engine on a test set of 453 colonoscopy reports and 226 pathology reports, considering accuracy in extracting the values of the target variables from text, and the reliability of the outcomes of the quality measures as computed from the NLP-extracted information. RESULTS: The average accuracy of the NLP engine over all variables was 0.89 (range: 0.62-1.0) and the average F measure over all variables was 0.74 (range: 0.49-0.89). The average agreement score, measured as Cohen's κ, between the manually established and NLP-derived outcomes of the quality measures was 0.62 (range: 0.09-0.86). DISCUSSION: For nine of the 19 colonoscopy quality measures, the agreement score was 0.70 or above, which we consider a sufficient score for the NLP-derived outcomes of these measures to be practically useful for quality measurement. CONCLUSION: The use of NLP for information extraction from free-text colonoscopy and pathology reports creates opportunities for large scale, routine quality measurement, which can support quality improvement in colonoscopy care. Henk Harkema, Wendy W. Chapman, Melissa I. Saul, Evan S. Dellon, Robert E. Schoen, Ateev Mehrotra |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Natural language processing: an introductionabstractOBJECTIVES: To provide an overview and tutorial of natural language processing (NLP) and modern NLP-system design. TARGET AUDIENCE: This tutorial targets the medical informatics generalist who has limited acquaintance with the principles behind NLP and/or limited knowledge of the current state of the art. SCOPE: We describe the historical evolution of NLP, and summarize common NLP sub-problems in this extensive field. We then provide a synopsis of selected highlights of medical NLP efforts. After providing a brief description of common machine-learning approaches that are being used for diverse NLP sub-problems, we discuss how modern NLP architectures are designed, with a summary of the Apache Foundation's Unstructured Information Management Architecture. We finally consider possible future directions for NLP, and reflect on the possible impact of IBM Watson on the medical field. Prakash M. Nadkarni, Lucila Ohno-Machado, Wendy W. Chapman |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | Anaphoric relations in the clinical narrative: corpus creationabstractOBJECTIVE: The long-term goal of this work is the automated discovery of anaphoric relations from the clinical narrative. The creation of a gold standard set from a cross-institutional corpus of clinical notes and high-level characteristics of that gold standard are described. METHODS: A standard methodology for annotation guideline development, gold standard annotations, and inter-annotator agreement (IAA) was used. RESULTS: The gold standard annotations resulted in 7214 markables, 5992 pairs, and 1304 chains. Each report averaged 40 anaphoric markables, 33 pairs, and seven chains. The overall IAA is high on the Mayo dataset (0.6607), and moderate on the University of Pittsburgh Medical Center (UPMC) dataset (0.4072). The IAA between each annotator and the gold standard is high (Mayo: 0.7669, 0.7697, and 0.9021; UPMC: 0.6753 and 0.7138). These results imply a quality corpus feasible for system development. They also suggest the complementary nature of the annotations performed by the experts and the importance of an annotator team with diverse knowledge backgrounds. LIMITATIONS: Only one of the annotators had the linguistic background necessary for annotation of the linguistic attributes. The overall generalizability of the guidelines will be further strengthened by annotations of data from additional sites. This will increase the overall corpus size and the representation of each relation type. CONCLUSION: The first step toward the development of an anaphoric relation resolver as part of a comprehensive natural language processing system geared specifically for the clinical narrative in the electronic medical record is described. The deidentified annotated corpus will be available to researchers. Guergana K. Savova, Wendy W. Chapman, Jiaping Zheng, Rebecca S. Jacobson |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Document-level classification of CT pulmonary angiography reports based on an extension of the ConText algorithm
Brian E. Chapman, Sean Lee, Hyunseok Peter Kang, Wendy W. Chapman |
J. Biomed. Informatics | 4 |
| 2011 | Coreference resolution: A review of general methodologies and applications in the clinical domain
Jiaping Zheng, Wendy W. Chapman, Rebecca S. Jacobson, Guergana K. Savova |
J. Biomed. Informatics | 2 |
| 2010 | Developing syndrome definitions based on consensus and current useabstractOBJECTIVE: Standardized surveillance syndromes do not exist but would facilitate sharing data among surveillance systems and comparing the accuracy of existing systems. The objective of this study was to create reference syndrome definitions from a consensus of investigators who currently have or are building syndromic surveillance systems. DESIGN: Clinical condition-syndrome pairs were catalogued for 10 surveillance systems across the United States and the representatives of these systems were brought together for a workshop to discuss consensus syndrome definitions. RESULTS: Consensus syndrome definitions were generated for the four syndromes monitored by the majority of the 10 participating surveillance systems: Respiratory, gastrointestinal, constitutional, and influenza-like illness (ILI). An important element in coming to consensus quickly was the development of a sensitive and specific definition for respiratory and gastrointestinal syndromes. After the workshop, the definitions were refined and supplemented with keywords and regular expressions, the keywords were mapped to standard vocabularies, and a web ontology language (OWL) ontology was created. LIMITATIONS: The consensus definitions have not yet been validated through implementation. CONCLUSION: The consensus definitions provide an explicit description of the current state-of-the-art syndromes used in automated surveillance, which can subsequently be systematically evaluated against real data to improve the definitions. The method for creating consensus definitions could be applied to other domains that have diverse existing definitions. Wendy W. Chapman, John N. Dowling, Atar Baer, David L. Buckeridge, Dennis Cochrane, Michael A. Conway, Peter L. Elkin, Jeremy U. Espino, Julia E. Gunn, Craig M. Hales, Lori Hutwagner, Mikaela Keller, Catherine Larson, Rebecca Noe, Anya Okhmatovskaia, Karen Olson, Marc Paladini, Matthew Scholer, Carol Sniegoski, William B. Lober |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Methodology to Develop and Evaluate a Semantic Representation for NLP
Jeannie Irwin, Henk Harkema, Lee M. Christensen, Titus Schleyer, Peter J. Haug, Wendy W. Chapman |
AMIA | 6 |
| 2009 | Developing a manually annotated clinical document corpus to identify phenotypic information for inflammatory bowel diseaseabstractBACKGROUND: Natural Language Processing (NLP) systems can be used for specific Information Extraction (IE) tasks such as extracting phenotypic data from the electronic medical record (EMR). These data are useful for translational research and are often found only in free text clinical notes. A key required step for IE is the manual annotation of clinical corpora and the creation of a reference standard for (1) training and validation tasks and (2) to focus and clarify NLP system requirements. These tasks are time consuming, expensive, and require considerable effort on the part of human reviewers. METHODS: Using a set of clinical documents from the VA EMR for a particular use case of interest we identify specific challenges and present several opportunities for annotation tasks. We demonstrate specific methods using an open source annotation tool, a customized annotation schema, and a corpus of clinical documents for patients known to have a diagnosis of Inflammatory Bowel Disease (IBD). We report clinician annotator agreement at the document, concept, and concept attribute level. We estimate concept yield in terms of annotated concepts within specific note sections and document types. RESULTS: Annotator agreement at the document level for documents that contained concepts of interest for IBD using estimated Kappa statistic (95% CI) was very high at 0.87 (0.82, 0.93). At the concept level, F-measure ranged from 0.61 to 0.83. However, agreement varied greatly at the specific concept attribute level. For this particular use case (IBD), clinical documents producing the highest concept yield per document included GI clinic notes and primary care notes. Within the various types of notes, the highest concept yield was in sections representing patient assessment and history of presenting illness. Ancillary service documents and family history and plan note sections produced the lowest concept yield. CONCLUSION: Challenges include defining and building appropriate annotation schemas, adequately training clinician annotators, and determining the appropriate level of information to be annotated. Opportunities include narrowing the focus of information extraction to use case specific note types and sections, especially in cases where NLP systems will be used to extract information from large repositories of electronic clinical note documents. Brett R. South, Shuying Shen, Makoto Jones, Jennifer H. Garvin, Matthew H. Samore, Wendy W. Chapman, Adi V. Gundlapalli |
BMC Bioinform. | 6 |
| 2009 | Current issues in biomedical text mining and natural language processing
Wendy W. Chapman, Kevin Cohen 0001 |
J. Biomed. Informatics | 1 |
| 2009 | What can natural language processing do for clinical decision support?
Dina Demner-Fushman, Wendy W. Chapman, Clement J. McDonald |
J. Biomed. Informatics | 2 |
| 2009 | ConText: An algorithm for determining negation, experiencer, and temporal status from clinical reports
Henk Harkema, John N. Dowling, Tyler Thornblade, Wendy W. Chapman |
J. Biomed. Informatics | 4 |
| 2008 | Identifying Data Sharing in Biomedical Literature
Heather A. Piwowar, Wendy W. Chapman |
AMIA | 2 |
| 2008 | Optimizing A Syndromic Surveillance Text Classifier for Influenza-like Illness: Does Document Source Matter?
Brett R. South, Wendy W. Chapman, Sylvain Delisle, Shuying Shen, Ericka Kalp, Trish Perl, Matthew H. Samore, Adi V. Gundlapalli |
AMIA | 2 |
| 2008 | Analysis of a Failed Clinical Decision Support System for Management of Congestive Heart Failure
Rajiv Wadhwa, Douglas B. Fridsma, Melissa I. Saul, Louis E. Penrod, Shyam Visweswaran, Gregory F. Cooper, Wendy W. Chapman |
AMIA | 7 |
| 2008 | Evaluation of preprocessing techniques for chief complaint classification
Jagan Dara, John N. Dowling, Debbie A. Travers, Gregory F. Cooper, Wendy W. Chapman |
J. Biomed. Informatics | 5 |
| 2007 | Methods Paper: Heuristic Sample Selection to Minimize Reference Standard Training Set for a Part-Of-Speech TaggerabstractPart-of-speech tagging represents an important first step for most medical natural language processing (NLP) systems. The majority of current statistically-based POS taggers are trained using a general English corpus. Consequently, these systems perform poorly on medical text. Annotated medical corpora are difficult to develop because of the time and labor required. We investigated a heuristic-based sample selection method to minimize annotated corpus size for retraining a Maximum Entropy (ME) POS tagger. We developed a manually annotated domain specific corpus (DSC) of surgical pathology reports and a domain specific lexicon (DL). We sampled the DSC using two heuristics to produce smaller training sets and compared the retrained performance against (1) the original ME modeled tagger trained on general English, (2) the ME tagger retrained on the DL, and (3) the MedPost tagger trained on MEDLINE abstracts. RESULTS showed that the ME tagger retrained with a DSC was superior to the tagger retrained with the DL, and also superior to MedPost. Heuristic methods for sample selection produced performance equivalent to use of the entire training set, but with many fewer sentences. Learning curve analysis showed that sample selection would enable an 84% decrease in the size of the training set without a decrement in performance. We conclude that heuristic sample selection can be used to markedly reduce human annotation requirements for training of medical NLP systems. Kaihong Liu, Wendy W. Chapman, Rebecca Hwa, Rebecca S. Jacobson |
J. Am. Medical Informatics Assoc. | 2 |
| 2006 | Evaluating the Effectiveness of Four Contextual Features in Classifying Annotated Clinical Conditions in Emergency Department Reports
David Chu, John N. Dowling, Wendy W. Chapman |
AMIA | 3 |
| 2006 | Inductive creation of an annotation schema for manually indexing clinical conditions from emergency department reports
Wendy W. Chapman, John N. Dowling |
J. Biomed. Informatics | 1 |
| 2005 | Automating Tissue Bank Annotation from Pathology Reports - Comparison to a Gold Standard Expert Annotation Set
Kaihong Liu, Kevin J. Mitchell, Wendy W. Chapman, Rebecca S. Jacobson |
AMIA | 3 |
| 2005 | A Software Tool to Assist Researchers in Coding Free Text Clinical Reports
Manoj Ramachandran, Gregory F. Cooper, Wendy W. Chapman, John N. Dowling |
AMIA | 3 |
| 2005 | Classifying free-text triage chief complaints into syndromic categories with natural language processing
Wendy W. Chapman, Lee M. Christensen, Michael M. Wagner 0001, Peter J. Haug, Oleg Ivanov, John N. Dowling, Robert T. Olszewski |
Artif. Intell. Medicine | 1 |
| 2005 | Research Paper: Generating a Reliable Reference Standard Set for Syndromic Case ClassificationabstractOBJECTIVE: To generate and measure the reliability for a reference standard set with representative cases from seven broad syndromic case definitions and several narrower syndromic definitions used for biosurveillance. DESIGN: From 527,228 eligible patients between 1990 and 2003, we generated a set of patients potentially positive for seven syndromes by classifying all eligible patients according to their ICD-9 primary discharge diagnoses. We selected a representative subset of the cases for chart review by physicians, who read emergency department reports and assigned values to 14 variables related to the seven syndromes. MEASUREMENTS: (1) Positive predictive value of the ICD-9 diagnoses; (2) prevalence of the syndromic definitions and related variables; (3) agreement between physician raters demonstrated by kappa, kappa corrected for bias and prevalence, and Finn's r; and (4) reliability of the reference standard classifications demonstrated by generalizability coefficients. RESULTS: Positive predictive value for ICD-9 classification ranged from 0.33 for botulinic to 0.86 for gastrointestinal. We generated between 80 and 566 positive cases for six of the seven syndromic definitions. Rash syndrome exhibited low prevalence (34 cases). Agreement between physician raters was high, with kappa > 0.70 for most variables. Ratings showed no bias. Finn's r was >0.70 for all variables. Generalizability coefficients were >0.70 for all variables but three. CONCLUSION: Of the 27 syndromes generated by the 14 variables, 21 showed high enough prevalence, agreement, and reliability to be used as reference standard definitions against which an automated syndromic classifier could be compared. Syndromic definitions that showed poor agreement or low prevalence include febrile botulinic syndrome, febrile and nonfebrile rash syndrome, respiratory syndrome explained by a nonrespiratory or noninfectious diagnosis, and febrile and nonfebrile gastrointestinal syndrome explained by a nongastrointestinal or noninfectious diagnosis. Wendy W. Chapman, John N. Dowling, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2004 | Fever detection from free-text clinical records for biosurveillance
Wendy W. Chapman, John N. Dowling, Michael M. Wagner 0001 |
J. Biomed. Informatics | 1 |
| 2003 | Research Paper: Creating a Text Classifier to Detect Radiology Reports Describing Mediastinal Findings Associated with Inhalational Anthrax and Other DisordersabstractOBJECTIVE: The aim of this study was to create a classifier for automatic detection of chest radiograph reports consistent with the mediastinal findings of inhalational anthrax. DESIGN: The authors used the Identify Patient Sets (IPS) system to create a key word classifier for detecting reports describing mediastinal findings consistent with anthrax and compared their performances on a test set of 79,032 chest radiograph reports. MEASUREMENTS: Area under the ROC curve was the main outcome measure of the IPS classifier. Sensitivity and specificity of an initial IPS model were calculated based on an existing key word search and were compared against a Boolean version of the IPS classifier. RESULTS: The IPS classifier received an area under the ROC curve of 0.677 (90% CI = 0.628 to 0.772) with a specificity of 0.99 and maximum sensitivity of 0.35. The initial IPS model attained a specificity of 1.0 and a sensitivity of 0.04. CONCLUSION: The IPS system is a useful tool for helping domain experts create a statistical key word classifier for textual reports that is a potentially useful component in surveillance of radiographic findings suspicious for anthrax. Wendy W. Chapman, Gregory F. Cooper, Paul Hanbury, Brian E. Chapman, Lee H. Harrison, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2003 | Application of Information Technology: Automated Syndromic Surveillance for the 2002 Winter OlympicsabstractThe 2002 Olympic Winter Games were held in Utah from February 8 to March 16, 2002. Following the terrorist attacks on September 11, 2001, and the anthrax release in October 2001, the need for bioterrorism surveillance during the Games was paramount. A team of informaticists and public health specialists from Utah and Pittsburgh implemented the Real-time Outbreak and Disease Surveillance (RODS) system in Utah for the Games in just seven weeks. The strategies and challenges of implementing such a system in such a short time are discussed. The motivation and cooperation inspired by the 2002 Olympic Winter Games were a powerful driver in overcoming the organizational issues. Over 114,000 acute care encounters were monitored between February 8 and March 31, 2002. No outbreaks of public health significance were detected. The system was implemented successfully and operational for the 2002 Olympic Winter Games and remains operational today. Per H. Gesteland, Reed M. Gardner, Fu-Chiang Tsui, Jeremy U. Espino, Robert T. Rolfs, Brent C. James, Wendy W. Chapman, Andrew W. Moore 0001, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2002 | In their own words? A terminological analysis of e-mail to a cancer information service
Catherine Arnott Smith, P. Zoë Stavri, Wendy W. Chapman |
AMIA | 3 |
| 2002 | Creating a Software Tool for the Clinical Researcher - the IPS System
Bruce G. Buchanan, Wendy W. Chapman, Gregory F. Cooper, Paul Hanbury, Mehmet Kayaalp 0002, Manoj Ramachandran, Melissa I. Saul |
AMIA | 2 |
| 2002 | Rapid deployment of an electronic disease surveillance system in the state of Utah for the 2002 Olympic Winter Games
Per H. Gesteland, Michael M. Wagner 0001, Wendy W. Chapman, Jeremy U. Espino, Fu-Chiang Tsui, Reed M. Gardner, Robert T. Rolfs, Virginia M. Dato, Brent C. James, Peter J. Haug |
AMIA | 3 |
| 2002 | Accuracy of three classifiers of acute gastrointestinal syndrome for syndromic surveillance
Oleg Ivanov, Michael M. Wagner 0001, Wendy W. Chapman, Robert T. Olszewski |
AMIA | 3 |
| 2002 | Data, network, and application: technical description of the Utah RODS Winter Olympic Biosurveillance System
Fu-Chiang Tsui, Jeremy U. Espino, Michael M. Wagner 0001, Per H. Gesteland, Oleg Ivanov, Robert T. Olszewski, Xiaoming Zeng, Wendy W. Chapman, Weng-Keen Wong, Andrew W. Moore 0001 |
AMIA | 9 |
| 2001 | Combining decision support methodologies to diagnose pneumonia
Dominik Aronsky, Marcelo Fiszman, Wendy W. Chapman, Peter J. Haug |
AMIA | 3 |
| 2001 | Evaluation of negation phrases in narrative clinical reports
Wendy W. Chapman, Will Bridewell, Paul Hanbury, Gregory F. Cooper, Bruce G. Buchanan |
AMIA | 1 |
| 2001 | IPS: A System That Uses Machine Learning to Help Locate Patient Records for Clinical Research
Gregory F. Cooper, Bruce G. Buchanan, Wendy W. Chapman, Paul Hanbury, Mehmet Kayaalp 0002, Melissa I. Saul |
AMIA | 3 |
| 2001 | A Simple Algorithm for Identifying Negated Findings and Diseases in Discharge Summaries
Wendy W. Chapman, Will Bridewell, Paul Hanbury, Gregory F. Cooper, Bruce G. Buchanan |
J. Biomed. Informatics | 1 |
| 2001 | A Comparison of Classification Algorithms to Automatically Identify Chest X-Ray Reports That Support Pneumonia
Wendy W. Chapman, Marcelo Fiszman, Brian E. Chapman, Peter J. Haug |
J. Biomed. Informatics | 1 |
| 2000 | Contribution of a speech recognition system to a computerized pneumonia guideline in the emergency department
Wendy W. Chapman, Dominik Aronsky, Marcelo Fiszman, Peter J. Haug |
AMIA | 1 |
| 2000 | Using Decision Tree Classifiers to Confirm Pneumonia Diagnosis
David D. Eardley, Dominik Aronsky, Wendy W. Chapman, Peter J. Haug |
AMIA | 3 |
| 2000 | Research Paper: Automatic Detection of Acute Bacterial Pneumonia from Chest X-ray ReportsabstractOBJECTIVE: To evaluate the performance of a natural language processing system in extracting pneumonia-related concepts from chest x-ray reports. DESIGN: Four physicians, three lay persons, a natural language processing system, and two keyword searches (designated AAKS and KS) detected the presence or absence of three pneumonia-related concepts and inferred the presence or absence of acute bacterial pneumonia from 292 chest x-ray reports. Gold standard: Majority vote of three independent physicians. Reliability of the gold standard was measured. OUTCOME MEASURES: Recall, precision, specificity, and agreement (using Finn's R: statistic) with respect to the gold standard. Differences between the physicians and the other subjects were tested using the McNemar test for each pneumonia concept and for the disease inference of acute bacterial pneumonia. RESULTS: Reliability of the reference standard ranged from 0.86 to 0.96. Recall, precision, specificity, and agreement (Finn R:) for the inference on acute bacterial pneumonia were, respectively, 0.94, 0.87, 0.91, and 0.84 for physicians; 0.95, 0.78, 0.85, and 0.75 for natural language processing system; 0.46, 0.89, 0.95, and 0.54 for lay persons; 0.79, 0.63, 0.71, and 0.49 for AAKS; and 0.87, 0.70, 0.77, and 0.62 for KS. The McNemar pairwise comparisons showed differences between one physician and the natural language processing system for the infiltrate concept and between another physician and the natural language processing system for the inference on acute bacterial pneumonia. The comparisons also showed that most physicians were significantly different from the other subjects in all pneumonia concepts and the disease inference. CONCLUSION: In extracting pneumonia related concepts from chest x-ray reports, the performance of the natural language processing system was similar to that of physicians and better than that of lay persons and keyword searches. The encoded pneumonia information has the potential to support several pneumonia-related applications used in our institution. The applications include a decision support system called the antibiotic assistant, a computerized clinical protocol for pneumonia, and a quality assurance application in the radiology department. Marcelo Fiszman, Wendy W. Chapman, Dominik Aronsky, R. Scott Evans, Peter J. Haug |
J. Am. Medical Informatics Assoc. | 2 |
| 1999 | Correct vs. Parsed Data for Inferring Pneumonia in Chest X-ray Reports
Wendy W. Chapman, Marcelo Fiszman, Peter J. Haug |
AMIA | 1 |
| 1999 | Comparing expert systems for identifying chest x-ray reports that support pneumonia
Wendy W. Chapman, Peter J. Haug |
AMIA | 1 |
| 1999 | Automatic identification of pneumonia related concepts on chest x-ray reports
Marcelo Fiszman, Wendy W. Chapman, R. Scott Evans, Peter J. Haug |
AMIA | 2 |
| 1998 | Bayesian modeling for linking causally related observations in chest X-ray reports
Wendy W. Chapman, Peter J. Haug |
AMIA | 1 |