Brett R. South

dblp:99/8246 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
4since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 9 first-author · 4 since 2021
YearPublicationVenuePosition
2022 Sampling Adverse Drug Events in Outpatient Clinical Notes for Natural Language Processing Tasks
Joseph M. Plasek, Abigail Salem, Stuart R. Lipsitz, Mary G. Amato, Dinah Foer, Heba Edrees, Suzanne V. Blackley, Brett R. South, Amol Rajmane, Mario Lorenzo, Paul Felt, Brendan Bull, Gretchen Purcell Jackson, Henry Feldman, David W. Bates, Li Zhou 0007
AMIA8
2021 Deploying Conversational Agents to Facilitate Housing Assistance Needs Resulting from COVID-19
Brett R. South, Anita M. Preininger, Piyush Parmar, Rubina F. Rizvi, David Brotman, Shira Alevy, Mollie McKillop, Gretchen Purcell Jackson, William Kassler
AMIA1
2021 Extraction of Ambiguous Phrases Found Within Adverse Drug Event Mentions Using A Natural Language Processing-Based Annotation tool
Zhoujun Sun, Rubina F. Rizvi, Shilo Anders, Brett R. South, Elisabeth Scheufele, Karlis Draulis, Henry J. Feldman
AMIA4
2021 Leveraging conversational technology to answer common COVID-19 questions
abstract
The rapidly evolving science about the Coronavirus Disease 2019 (COVID-19) pandemic created unprecedented health information needs and dramatic changes in policies globally. We describe a platform, Watson Assistant (WA), which has been used to develop conversational agents to deliver COVID-19 related information. We characterized the diverse use cases and implementations during the early pandemic and measured adoption through a number of users, messages sent, and conversational turns (ie, pairs of interactions between users and agents). Thirty-seven institutions in 9 countries deployed COVID-19 conversational agents with WA between March 30 and August 10, 2020, including 24 governmental agencies, 7 employers, 5 provider organizations, and 1 health plan. Over 6.8 million messages were delivered through the platform. The mean number of conversational turns per session ranged between 1.9 and 3.5. Our experience demonstrates that conversational technologies can be rapidly deployed for pandemic response and are adopted globally by a wide range of users.
Mollie McKillop, Brett R. South, Anita M. Preininger, Mitch Mason, Gretchen Purcell Jackson
J. Am. Medical Informatics Assoc.2
2020 Identifying and Leveraging Public Data Sources with Structured Social Determinants of Health Information for Observational Health Research
Irene Dankwa-Mullan, Mollie McKillop, Metasebya Solomon, Anita M. Preininger, Mark C. Roebuck, Yull Arriaga, Judy George, Gretchen Purcell Jackson, Brett R. South
AMIA9
2020 Development and Training of an Artificial Intelligence (AI)-Based Clinical Trial Matching Tool for Oncology
Tufia C. Haddad, Konstantinos Levantakos, Brett R. South, Dilhan Weeraratne, S. John Weroha, Andrea E. Hendrickson, Amit Mahipal, Megan Sands-Lincoln, Nawshin Kutub, Melissa Rammage, Eric Will, Jane Helgeson, Jane L. Snowdon, Gretchen Purcell Jackson
AMIA3
2020 A Modular Editorial Content Curation Framework for Pharmacological and Patient Knowledge Management
Elisabeth Scheufele, Jason Hatanaka, Mya Baca, Stacy Laclaire, Jeff Heiland, Nawshin Kutub, Brett R. South, Gretchen Purcell Jackson
AMIA7
2019 Developing Synthetic VA Healthcare Data in OMOP CDM Model
Jiantao Bian, Hamid Saoudian, Brett R. South, Kristine E. Lynch, Benjamin Viernes, Michael E. Matheny, Scott L. DuVall
AMIA3
2018 The Department of Defense (DoD) and Department of Veterans Affairs (VA) Infrastructure for Clinical Intelligence (DaVINCI)
Scott L. DuVall, Michael E. Matheny, Ildar R. Ibragimov, Trey D. Oats, Jay N. Tucker, Brett R. South, Augie Turano, Hamid Saoudian, Casey Kangas, Keith D. Hofmann, Wendy Funk, Chris Nichols, Albert Bonnema, Louis Ferrucci, Jonathan R. Nebeker
AMIA6
2018 Detecting Current Episodes of Cholecystitis-related Pain from Veterans Affairs Clinical Notes using Natural Language Processing
Danielle L. Mowery, Luke Martin, Brett R. South, Ellen Morrow, William Peche, Eric Wiesner, Wendy W. Chapman, Benjamin S. Brooke
AMIA3
2018 Annotating Social Determinants of Health and Functional Status Information Using Publicly Accessible Corpora
Ruth M. Reeves, Brett R. South, Glenn T. Gobbel, Lee M. Christensen, Wendy W. Chapman, Michael E. Matheny, Jeremiah R. Brown
AMIA2
2018 Flipping the Model for Biomedical Informatics Research
Brett R. South, Kristine W. Lynch, Michael E. Matheny, Julie A. Lynch, Olga Efimova, Catherine Chanfreau-Coffinier, Olga V. Patterson, Benjamin Viernes, Scott L. DuVall
AMIA1
2016 Automated Extraction of Disease Activity Component Measures From Electronic Medical Records Using Natural Language Processing
Brett R. South, Shobhit Mehortra, Jianwei Leng, Sophia Lu, Brian C. Sauer, Grant Cannon
AMIA1
2015 The Informatics Sculptor & the Clinical Annotator: Effective Annotation Strategies
Nancy Gentry, Elizabeth Hanchrow, Glenn T. Gobbel, Brett R. South, Steven M. Bradley, Ruth M. Reeves
AMIA4
2015 RapTAT: A Tool for Assisted Annotation and Reviewer Training via Online Machine Learning
Glenn T. Gobbel, Ruth M. Reeves, Brett R. South, Wendy W. Chapman, Jay Jarman, Steven Lay, Michael E. Matheny
AMIA3
2015 Annotating ADLs and IADLs in Veterans Affairs Clinical Documents
Brett R. South, Danielle L. Mowery, Lee M. Christensen, Adi V. Gundlapalli, Melissa Tharp, Marzieh Vali, Marjorie Carter, Mike Conway, Salomeh Keyhani, Wendy W. Chapman
AMIA1
2015 Evaluating the state of the art in disorder recognition and normalization of the clinical narrative
abstract
OBJECTIVE: The ShARe/CLEF eHealth 2013 Evaluation Lab Task 1 was organized to evaluate the state of the art on the clinical text in (i) disorder mention identification/recognition based on Unified Medical Language System (UMLS) definition (Task 1a) and (ii) disorder mention normalization to an ontology (Task 1b). Such a community evaluation has not been previously executed. Task 1a included a total of 22 system submissions, and Task 1b included 17. Most of the systems employed a combination of rules and machine learners. MATERIALS AND METHODS: We used a subset of the Shared Annotated Resources (ShARe) corpus of annotated clinical text--199 clinical notes for training and 99 for testing (roughly 180 K words in total). We provided the community with the annotated gold standard training documents to build systems to identify and normalize disorder mentions. The systems were tested on a held-out gold standard test set to measure their performance. RESULTS: For Task 1a, the best-performing system achieved an F1 score of 0.75 (0.80 precision; 0.71 recall). For Task 1b, another system performed best with an accuracy of 0.59. DISCUSSION: Most of the participating systems used a hybrid approach by supplementing machine-learning algorithms with features generated by rules and gazetteers created from the training data and from external resources. CONCLUSIONS: The task of disorder normalization is more challenging than that of identification. The ShARe corpus is available to the community as a reference standard for future studies.
Sameer Pradhan, Noémie Elhadad, Brett R. South, David Martínez 0001, Lee M. Christensen, Amy Vogel, Hanna Suominen, Wendy W. Chapman, Guergana K. Savova
J. Am. Medical Informatics Assoc.3
2014 Extracting Concepts Related to a Homelessness from the Free Text of VA Electronic Medical Records
Adi V. Gundlapalli, Marjorie Carter, Guy Divita, Shuying Shen, Miland N. Palmer, Brett R. South, Begum Durgahee, Matthew H. Samore
AMIA6
2014 A System Usability Study Assessing a Machine-Assisted Interactive Interface to Support Annotation of Protected Health Information in Clinical Texts
Brett R. South, Danielle L. Mowery, Chris Leng, Stéphane M. Meystre, Wendy W. Chapman
AMIA1
2014 Text de-identification for privacy protection: A study of its impact on clinical text information content
Stéphane M. Meystre, Óscar Ferrández, F. Jeffrey Friedlin, Brett R. South, Shuying Shen, Matthew H. Samore
J. Biomed. Informatics4
2014 Evaluating the effects of machine pre-annotation and an interactive annotation interface on manual de-identification of clinical text
abstract
The Health Insurance Portability and Accountability Act (HIPAA) Safe Harbor method requires removal of 18 types of protected health information (PHI) from clinical documents to be considered "de-identified" prior to use for research purposes. Human review of PHI elements from a large corpus of clinical documents can be tedious and error-prone. Indeed, multiple annotators may be required to consistently redact information that represents each PHI class. Automated de-identification has the potential to improve annotation quality and reduce annotation time. For instance, using machine-assisted annotation by combining de-identification system outputs used as pre-annotations and an interactive annotation interface to provide annotators with PHI annotations for "curation" rather than manual annotation from "scratch" on raw clinical documents. In order to assess whether machine-assisted annotation improves the reliability and accuracy of the reference standard quality and reduces annotation effort, we conducted an annotation experiment. In this annotation study, we assessed the generalizability of the VA Consortium for Healthcare Informatics Research (CHIR) annotation schema and guidelines applied to a corpus of publicly available clinical documents called MTSamples. Specifically, our goals were to (1) characterize a heterogeneous corpus of clinical documents manually annotated for risk-ranked PHI and other annotation types (clinical eponyms and person relations), (2) evaluate how well annotators apply the CHIR schema to the heterogeneous corpus, (3) compare whether machine-assisted annotation (experiment) improves annotation quality and reduces annotation time compared to manual annotation (control), and (4) assess the change in quality of reference standard coverage with each added annotator's annotations.
Brett R. South, Danielle L. Mowery, Ying Suo, Jianwei Leng, Óscar Ferrández, Stéphane M. Meystre, Wendy W. Chapman
J. Biomed. Informatics1
2013 Using Natural Language Processing on the Free Text of Clinical Documents to Screen for Evidence of Homelessness Among US Veterans
Adi V. Gundlapalli, Marjorie Carter, Miland N. Palmer, Thomas Ginter, Andrew Redd, Steve Pickard, Shuying Shen, Brett R. South, Guy Divita, Scott L. DuVall, Thien M. Nguyen, Leonard W. D'Avolio, Matthew H. Samore
AMIA8
2013 Creating a Reference Standard of Acronym and Abbreviation Annotations for the ShARe/CLEF eHealth Challenge 2013
Danielle L. Mowery, Brett R. South, Jianwei Leng, Laura-Maria Peltonen, Riitta Danielsson-Ojala, Sanna Salanterä, Wendy W. Chapman
AMIA2
2013 BoB, a best-of-breed automated text de-identification system for VHA clinical documents
abstract
OBJECTIVE: De-identification allows faster and more collaborative clinical research while protecting patient confidentiality. Clinical narrative de-identification is a tedious process that can be alleviated by automated natural language processing methods. The goal of this research is the development of an automated text de-identification system for Veterans Health Administration (VHA) clinical documents. MATERIALS AND METHODS: We devised a novel stepwise hybrid approach designed to improve the current strategies used for text de-identification. The proposed system is based on a previous study on the best de-identification methods for VHA documents. This best-of-breed automated clinical text de-identification system (aka BoB) tackles the problem as two separate tasks: (1) maximize patient confidentiality by redacting as much protected health information (PHI) as possible; and (2) leave de-identified documents in a usable state preserving as much clinical information as possible. RESULTS: We evaluated BoB with a manually annotated corpus of a variety of VHA clinical notes, as well as with the 2006 i2b2 de-identification challenge corpus. We present evaluations at the instance- and token-level, with detailed results for BoB's main components. Moreover, an existing text de-identification system was also included in our evaluation. DISCUSSION: BoB's design efficiently takes advantage of the methods implemented in its pipeline, resulting in high sensitivity values (especially for sensitive PHI categories) and a limited number of false positives. CONCLUSIONS: Our system successfully addressed VHA clinical document de-identification, and its hybrid stepwise design demonstrates robustness and efficiency, prioritizing patient confidentiality while leaving most clinical information intact.
Óscar Ferrández, Brett R. South, Shuying Shen, F. Jeffrey Friedlin, Matthew H. Samore, Stéphane M. Meystre
J. Am. Medical Informatics Assoc.2
2012 Generalizability and Comparison of Automatic Clinical Text De-Identification Methods and Resources
Óscar Ferrández, Brett R. South, Shuying Shen, F. Jeffrey Friedlin, Matthew H. Samore, Stéphane M. Meystre
AMIA2
2012 CASPR: Friendly Annotation Management
Tyler Forbush, Brad Adams, Shuying Shen, Brett R. South, Jonathan R. Nebeker, Scott L. DuVall
AMIA4
2012 A Survey of VHA Privacy Officers for the External Use of Automatically De-Identified Clinical Documents
Neil Nokes, Stéphane M. Meystre, Brett R. South, Jeffrey Scehnet, Shuying Shen, Óscar Ferrández, F. Jeffrey Friedlin, Matthew Maw, Matthew H. Samore
AMIA3
2012 The Relationship Between Structural Characteristics of Electronic Clinical Texts and Ratings of Document Quality
Shuying Shen, Brett R. South, Jorie Butler, Robyn Barrus, Charlene R. Weir
AMIA2
2012 On the Road Towards Developing a Publicly Available Corpus of De-identified Clinical Texts
Brett R. South, Danielle L. Mowery, Óscar Ferrández, Shuying Shen, Ying Suo, Annie Chen, Stéphane M. Meystre, Wendy W. Chapman
AMIA1
2012 Evaluation and Visualization of Human Annotator Learning Patterns
Ying Suo, Shuying Shen, Scott L. DuVall, Özlem Uzuner, Brett R. South
AMIA5
2012 Automated extraction of ejection fraction for quality measurement using regular expressions in Unstructured Information Management Architecture (UIMA) for heart failure
abstract
OBJECTIVES: Left ventricular ejection fraction (EF) is a key component of heart failure quality measures used within the Department of Veteran Affairs (VA). Our goals were to build a natural language processing system to extract the EF from free-text echocardiogram reports to automate measurement reporting and to validate the accuracy of the system using a comparison reference standard developed through human review. This project was a Translational Use Case Project within the VA Consortium for Healthcare Informatics. MATERIALS AND METHODS: We created a set of regular expressions and rules to capture the EF using a random sample of 765 echocardiograms from seven VA medical centers. The documents were randomly assigned to two sets: a set of 275 used for training and a second set of 490 used for testing and validation. To establish the reference standard, two independent reviewers annotated all documents in both sets; a third reviewer adjudicated disagreements. RESULTS: System test results for document-level classification of EF of <40% had a sensitivity (recall) of 98.41%, a specificity of 100%, a positive predictive value (precision) of 100%, and an F measure of 99.2%. System test results at the concept level had a sensitivity of 88.9% (95% CI 87.7% to 90.0%), a positive predictive value of 95% (95% CI 94.2% to 95.9%), and an F measure of 91.9% (95% CI 91.2% to 92.7%). DISCUSSION: An EF value of <40% can be accurately identified in VA echocardiogram reports. CONCLUSIONS: An automated information extraction system can be used to accurately extract EF for quality measurement.
Jennifer H. Garvin, Scott L. DuVall, Brett R. South, Bruce E. Bray, Daniel Bolton, Julia Heavirland, Steve Pickard, Paul Heidenreich, Shuying Shen, Charlene R. Weir, Matthew H. Samore, Mary K. Goldstein
J. Am. Medical Informatics Assoc.3
2012 Evaluating the state of the art in coreference resolution for electronic medical records
abstract
BACKGROUND: The fifth i2b2/VA Workshop on Natural Language Processing Challenges for Clinical Records conducted a systematic review on resolution of noun phrase coreference in medical records. Informatics for Integrating Biology and the Bedside (i2b2) and the Veterans Affair (VA) Consortium for Healthcare Informatics Research (CHIR) partnered to organize the coreference challenge. They provided the research community with two corpora of medical records for the development and evaluation of the coreference resolution systems. These corpora contained various record types (ie, discharge summaries, pathology reports) from multiple institutions. METHODS: The coreference challenge provided the community with two annotated ground truth corpora and evaluated systems on coreference resolution in two ways: first, it evaluated systems for their ability to identify mentions of concepts and to link together those mentions. Second, it evaluated the ability of the systems to link together ground truth mentions that refer to the same entity. Twenty teams representing 29 organizations and nine countries participated in the coreference challenge. RESULTS: The teams' system submissions showed that machine-learning and rule-based approaches worked best when augmented with external knowledge sources and coreference clues extracted from document structure. The systems performed better in coreference resolution when provided with ground truth mentions. Overall, the systems struggled in solving coreference resolution for cases that required domain knowledge.
Özlem Uzuner, Andreea Bodnari, Shuying Shen, Tyler Forbush, John Pestian, Brett R. South
J. Am. Medical Informatics Assoc.6
2011 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text
abstract
The 2010 i2b2/VA Workshop on Natural Language Processing Challenges for Clinical Records presented three tasks: a concept extraction task focused on the extraction of medical concepts from patient reports; an assertion classification task focused on assigning assertion types for medical problem concepts; and a relation classification task focused on assigning relation types that hold between medical problems, tests, and treatments. i2b2 and the VA provided an annotated reference standard corpus for the three tasks. Using this reference standard, 22 systems were developed for concept extraction, 21 for assertion classification, and 16 for relation classification. These systems showed that machine learning approaches could be augmented with rule-based systems to determine concepts, assertions, and relations. Depending on the task, the rule-based systems can either provide input for machine learning or post-process the output of machine learning. Ensembles of classifiers, information from unlabeled data, and external knowledge sources can help when the training data are inadequate.
Özlem Uzuner, Brett R. South, Shuying Shen, Scott L. DuVall
J. Am. Medical Informatics Assoc.2
2010 Textractor: a hybrid system for medications and reason for their prescription extraction from clinical text documents
abstract
UNLABELLED: OBJECTIVE To describe a new medication information extraction system-Textractor-developed for the 'i2b2 medication extraction challenge'. The development, functionalities, and official evaluation of the system are detailed. DESIGN: Textractor is based on the Apache Unstructured Information Management Architecture (UMIA) framework, and uses methods that are a hybrid between machine learning and pattern matching. Two modules in the system are based on machine learning algorithms, while other modules use regular expressions, rules, and dictionaries, and one module embeds MetaMap Transfer. MEASUREMENTS: The official evaluation was based on a reference standard of 251 discharge summaries annotated by all teams participating in the challenge. The metrics used were recall, precision, and the F(1)-measure. They were calculated with exact and inexact matches, and were averaged at the level of systems and documents. RESULTS: The reference metric for this challenge, the system-level overall F(1)-measure, reached about 77% for exact matches, with a recall of 72% and a precision of 83%. Performance was the best with route information (F(1)-measure about 86%), and was good for dosage and frequency information, with F(1)-measures of about 82-85%. Results were not as good for durations, with F(1)-measures of 36-39%, and for reasons, with F(1)-measures of 24-27%. CONCLUSION: The official evaluation of Textractor for the i2b2 medication extraction challenge demonstrated satisfactory performance. This system was among the 10 best performing systems in this challenge.
Stéphane M. Meystre, Julien Thibault, Shuying Shen, John F. Hurdle, Brett R. South
J. Am. Medical Informatics Assoc.5
2009 Inductive Creation of an Annotation Schema and a Reference Standard for De-identification of VA Electronic Clinical Notes
Jeanmarie Mayer, Shuying Shen, Brett R. South, Stéphane M. Meystre, F. Jeffrey Friedlin, William R. Ray, Matthew H. Samore
AMIA3
2009 Developing a manually annotated clinical document corpus to identify phenotypic information for inflammatory bowel disease
abstract
BACKGROUND: Natural Language Processing (NLP) systems can be used for specific Information Extraction (IE) tasks such as extracting phenotypic data from the electronic medical record (EMR). These data are useful for translational research and are often found only in free text clinical notes. A key required step for IE is the manual annotation of clinical corpora and the creation of a reference standard for (1) training and validation tasks and (2) to focus and clarify NLP system requirements. These tasks are time consuming, expensive, and require considerable effort on the part of human reviewers. METHODS: Using a set of clinical documents from the VA EMR for a particular use case of interest we identify specific challenges and present several opportunities for annotation tasks. We demonstrate specific methods using an open source annotation tool, a customized annotation schema, and a corpus of clinical documents for patients known to have a diagnosis of Inflammatory Bowel Disease (IBD). We report clinician annotator agreement at the document, concept, and concept attribute level. We estimate concept yield in terms of annotated concepts within specific note sections and document types. RESULTS: Annotator agreement at the document level for documents that contained concepts of interest for IBD using estimated Kappa statistic (95% CI) was very high at 0.87 (0.82, 0.93). At the concept level, F-measure ranged from 0.61 to 0.83. However, agreement varied greatly at the specific concept attribute level. For this particular use case (IBD), clinical documents producing the highest concept yield per document included GI clinic notes and primary care notes. Within the various types of notes, the highest concept yield was in sections representing patient assessment and history of presenting illness. Ancillary service documents and family history and plan note sections produced the lowest concept yield. CONCLUSION: Challenges include defining and building appropriate annotation schemas, adequately training clinician annotators, and determining the appropriate level of information to be annotated. Opportunities include narrowing the focus of information extraction to use case specific note types and sections, especially in cases where NLP systems will be used to extract information from large repositories of electronic clinical note documents.
Brett R. South, Shuying Shen, Makoto Jones, Jennifer H. Garvin, Matthew H. Samore, Wendy W. Chapman, Adi V. Gundlapalli
BMC Bioinform.1
2008 Optimizing A Syndromic Surveillance Text Classifier for Influenza-like Illness: Does Document Source Matter?
Brett R. South, Wendy W. Chapman, Sylvain Delisle, Shuying Shen, Ericka Kalp, Trish Perl, Matthew H. Samore, Adi V. Gundlapalli
AMIA1