Stéphane M. Meystre

dblp:88/3487 · DBLP profile ↗
← Back
56ranked-venue papers
21as first author
9since 2021 · last 2024
0000-0002-7632-9625ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 56 · 21 first-author · 9 since 2021
YearPublicationVenuePosition
2024 Large language models for biomedicine: foundations, opportunities, challenges, and best practices
abstract
OBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications.
Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang
J. Am. Medical Informatics Assoc.8
2022 Post-Hoc Ensemble Generation for Clinical NLP: A Study of Concept Recognition, Normalization, and Context Attributes
Paul M. Heider, Ronak Pipaliya, Stéphane M. Meystre
AMIA4
2022 Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborative
abstract
OBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require.
Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai
J. Am. Medical Informatics Assoc.65
2021 Colonizing Microbiome as a Determinant of COVID-19 Outcome: A Pilot Study
Alexander V. Alekseyenko, Bashir Hamidi, Stéphane M. Meystre
AMIA3
2021 Overview and Descriptive Analysis of a New Ontology for Normalizing Section Types in Unstructured Clinical Notes
Paul M. Heider, Stéphane M. Meystre
AMIA2
2021 Evaluating the Downstream Performance Impact of Various Common Off-the-Shelf Clinical NLP Components
Paul M. Heider, Stéphane M. Meystre
AMIA2
2021 Clinical Concept Extraction Using Contextual String Embeddings
Paul M. Heider, Stéphane M. Meystre
AMIA3
2021 Natural Language Processing and COVID-19 Predictive Analytics to Enable and Optimize SARS-CoV-2 Pooled Testing
Stéphane M. Meystre, Paul M. Heider, Jihad S. Obeid, Alexander V. Alekseyenko, James E. Madory
AMIA1
2021 Natural language processing enabling COVID-19 predictive analytics to support data-driven patient advising and pooled testing
abstract
OBJECTIVE: The COVID-19 (coronavirus disease 2019) pandemic response at the Medical University of South Carolina included virtual care visits for patients with suspected severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection. The telehealth system used for these visits only exports a text note to integrate with the electronic health record, but structured and coded information about COVID-19 (eg, exposure, risk factors, symptoms) was needed to support clinical care and early research as well as predictive analytics for data-driven patient advising and pooled testing. MATERIALS AND METHODS: To capture COVID-19 information from multiple sources, a new data mart and a new natural language processing (NLP) application prototype were developed. The NLP application combined reused components with dictionaries and rules crafted by domain experts. It was deployed as a Web service for hourly processing of new data from patients assessed or treated for COVID-19. The extracted information was then used to develop algorithms predicting SARS-CoV-2 diagnostic test results based on symptoms and exposure information. RESULTS: The dedicated data mart and NLP application were developed and deployed in a mere 10-day sprint in March 2020. The NLP application was evaluated with good accuracy (85.8% recall and 81.5% precision). The SARS-CoV-2 testing predictive analytics algorithms were configured to provide patients with data-driven COVID-19 testing advices with a sensitivity of 81% to 92% and to enable pooled testing with a negative predictive value of 90% to 91%, reducing the required tests to about 63%. CONCLUSIONS: SARS-CoV-2 testing predictive analytics and NLP successfully enabled data-driven patient advising and pooled testing.
Stéphane M. Meystre, Paul M. Heider, Jihad S. Obeid, James E. Madory, Alexander V. Alekseyenko
J. Am. Medical Informatics Assoc.1
2020 A Meta-Analysis of Medical Concept Normalization Using Hierarchical Ontological Relations and Semantic Types
Paul M. Heider, Stéphane M. Meystre
AMIA3
2020 Comparative Study of Various Approaches for Ensemble-based De-identification of Electronic Health Record Narratives
Paul M. Heider, Stéphane M. Meystre
AMIA3
2020 Improving De-identification of Clinical Text with Contextualized Embeddings
Stéphane M. Meystre
AMIA2
2020 De-Identification of Clinical Text: Stakeholders' Perspectives and Acceptance of Automatic De-Identification
Stéphane M. Meystre, Jonathan C. Silverstein, Guergana K. Savova, Valentina Petkov, Bradley A. Malin
AMIA1
2020 Leveraging health system telehealth and informatics infrastructure to create a continuum of services for COVID-19 screening, testing, and treatment
abstract
OBJECTIVES: We describe our approach in using health information technology to provide a continuum of services during the coronavirus disease 2019 (COVID-19) pandemic. COVID-19 challenges and needs required health systems to rapidly redesign the delivery of care. MATERIALS AND METHODS: Our health system deployed 4 COVID-19 telehealth programs and 4 biomedical informatics innovations to screen and care for COVID-19 patients. Using programmatic and electronic health record data, we describe the implementation and initial utilization. RESULTS: Through collaboration across multidisciplinary teams and strategic planning, 4 telehealth program initiatives have been deployed in response to COVID-19: virtual urgent care screening, remote patient monitoring for COVID-19-positive patients, continuous virtual monitoring to reduce workforce risk and utilization of personal protective equipment, and the transition of outpatient care to telehealth. Biomedical informatics was integral to our institutional response in supporting clinical care through new and reconfigured technologies. Through linking the telehealth systems and the electronic health record, we have the ability to monitor and track patients through a continuum of COVID-19 services. DISCUSSION: COVID-19 has facilitated the rapid expansion and utilization of telehealth and health informatics services. We anticipate that patients and providers will view enhanced telehealth services as an essential aspect of the healthcare system. Continuation of telehealth payment models at the federal and private levels will be a key factor in whether this new uptake is sustained. CONCLUSIONS: There are substantial benefits in utilizing telehealth during the COVID-19, including the ability to rapidly scale the number of patients being screened and providing continuity of care.
Dee W. Ford, Jillian B. Harvey, James T. McElligott, Kathryn King, Kit N. Simpson, Shawn Valenta, Emily H. Warr, Tasia Walsh, Ellen Debenham, Carla Teasdale, Stéphane M. Meystre, Jihad S. Obeid, Christopher Metts, Leslie Lenert
J. Am. Medical Informatics Assoc.11
2020 Ensemble method-based extraction of medication and related information from clinical texts
abstract
OBJECTIVE: Accurate and complete information about medications and related information is crucial for effective clinical decision support and precise health care. Recognition and reduction of adverse drug events is also central to effective patient care. The goal of this research is the development of a natural language processing (NLP) system to automatically extract medication and adverse drug event information from electronic health records. This effort was part of the 2018 n2c2 shared task on adverse drug events and medication extraction. MATERIALS AND METHODS: The new NLP system implements a stacked generalization based on a search-based structured prediction algorithm for concept extraction. We trained 4 sequential classifiers using a variety of structured learning algorithms. To enhance accuracy, we created a stacked ensemble consisting of these concept extraction models trained on the shared task training data. We implemented a support vector machine model to identify related concepts. RESULTS: Experiments with the official test set showed that our stacked ensemble achieved an F1 score of 92.66%. The relation extraction model with given concepts reached a 93.59% F1 score. Our end-to-end system yielded overall micro-averaged recall, precision, and F1 score of 92.52%, 81.88% and 86.88%, respectively. Our NLP system for adverse drug events and medication extraction ranked within the top 5 of teams participating in the challenge. CONCLUSION: This study demonstrated that a stacked ensemble with a search-based structured prediction algorithm achieved good performance by effectively integrating the output of individual classifiers and could provide a valid solution for other clinical concept extraction tasks.
Stéphane M. Meystre
J. Am. Medical Informatics Assoc.2
2020 An artificial intelligence approach to COVID-19 infection risk assessment in virtual visits: A case report
abstract
OBJECTIVE: In an effort to improve the efficiency of computer algorithms applied to screening for coronavirus disease 2019 (COVID-19) testing, we used natural language processing and artificial intelligence-based methods with unstructured patient data collected through telehealth visits. MATERIALS AND METHODS: After segmenting and parsing documents, we conducted analysis of overrepresented words in patient symptoms. We then developed a word embedding-based convolutional neural network for predicting COVID-19 test results based on patients' self-reported symptoms. RESULTS: Text analytics revealed that concepts such as smell and taste were more prevalent than expected in patients testing positive. As a result, screening algorithms were adapted to include these symptoms. The deep learning model yielded an area under the receiver-operating characteristic curve of 0.729 for predicting positive results and was subsequently applied to prioritize testing appointment scheduling. CONCLUSIONS: Informatics tools such as natural language processing and artificial intelligence methods can have significant clinical impacts when applied to data streams early in the development of clinical systems for outbreak response.
Jihad S. Obeid, Stéphane M. Meystre, Paul M. Heider, Edward C. O'Bryan, Leslie Lenert
J. Am. Medical Informatics Assoc.4
2019 Cancer Type Classification by Jointly Using Words and Concepts from Electronic Health Record Text Notes
Stéphane M. Meystre
AMIA2
2019 Text De-Identification Impact on Subsequent Machine Learning Applications
Gary Underwood, Andrew Trice, Jean-Karlo Accetta, Stéphane M. Meystre
AMIA5
2018 Ensemble-based Methods to Improve De-identification of Electronic Health Record Narratives
Paul M. Heider, Stéphane M. Meystre
AMIA3
2018 Automatic Text De-Identification: How and When is it Acceptable?
Stéphane M. Meystre, David Carrell, Lynette Hirschman, John S. Aberdeen, Paul Fearn, Valentina Petkov, Jonathan C. Silverstein
AMIA1
2018 Clinical Text Automatic De-Identification to Support Large Scale Data Reuse and Sharing: Pilot Results
Stéphane M. Meystre, Paul M. Heider, Andrew Trice, Gary Underwood
AMIA1
2017 Semi-automated Ontology Development and Management System Applied to Medically Unexplained Syndromes in the U.S. Veterans Population
Stéphane M. Meystre, Kristina Doing-Harris
AIME1
2017 Exploiting Unlabeled Texts with Clustering-based Instance Selection for Medical Relation Classification
Ellen Riloff, Stéphane M. Meystre
AMIA3
2017 Identifying Falls Risk Screenings Not Documented with Administrative Codes Using Natural Language Processing
Vivienne J. Zhu, Tina Walker, Robert W. Warren, Peggy Jenny, Stéphane M. Meystre, Leslie Lenert
AMIA5
2017 Congestive heart failure information extraction framework for automated treatment performance measures assessment
abstract
OBJECTIVE: This paper describes a new congestive heart failure (CHF) treatment performance measure information extraction system - CHIEF - developed as part of the Automated Data Acquisition for Heart Failure project, a Veterans Health Administration project aiming at improving the detection of patients not receiving recommended care for CHF. DESIGN: CHIEF is based on the Apache Unstructured Information Management Architecture framework, and uses a combination of rules, dictionaries, and machine learning methods to extract left ventricular function mentions and values, CHF medications, and documented reasons for a patient not receiving these medications. MEASUREMENTS: The training and evaluation of CHIEF were based on subsets of a reference standard of various clinical notes from 1083 Veterans Health Administration patients. Domain experts manually annotated these notes to create our reference standard. Metrics used included recall, precision, and the F 1 -measure. RESULTS: In general, CHIEF extracted CHF medications with high recall (>0.990) and good precision (0.960-0.978). Mentions of Left Ventricular Ejection Fraction were also extracted with high recall (0.978-0.986) and precision (0.986-0.994), and quantitative values of Left Ventricular Ejection Fraction were found with 0.910-0.945 recall and with high precision (0.939-0.976). Reasons for not prescribing CHF medications were more difficult to extract, only reaching fair accuracy with about 0.310-0.400 recall and 0.250-0.320 precision. CONCLUSION: This study demonstrated that applying natural language processing to unlock the rich and detailed clinical information found in clinical narrative text notes makes fast and scalable quality improvement approaches possible, eventually improving management and outpatient treatment of patients suffering from CHF.
Stéphane M. Meystre, Glenn T. Gobbel, Michael E. Matheny, Andrew Redd, Bruce E. Bray, Jennifer H. Garvin
J. Am. Medical Informatics Assoc.1
2017 Extraction of left ventricular ejection fraction information from various types of clinical reports
Jennifer H. Garvin, Mary K. Goldstein, Tammy S. Hwang, Andrew Redd, Daniel Bolton, Paul Heidenreich, Stéphane M. Meystre
J. Biomed. Informatics8
2016 Temporally Classifying Clinical Events Relative to Document Creation Time
Abdulrahman Khalifa, Stéphane M. Meystre
AMIA2
2016 Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001
AMIA1
2016 Automated Dynamic Problem and Allergy Lists for Efficient Electronic Health Record Management
Stéphane M. Meystre, Jianyin Shao, Greg M. Jones
AMIA1
2015 Improving Detection of Reasons Not to Take a Medication by Leveraging Medication Prescription Status
Jennifer H. Garvin, Julia Heavirland, Stéphane M. Meystre
AMIA4
2015 State of the Art of Clinical Narrative Report De-Identification and Its Future
Özlem Uzuner, John S. Aberdeen, Stéphane M. Meystre, Mehmet Kayaalp 0002
AMIA3
2015 Adapting existing natural language processing resources for cardiovascular risk factors identification in clinical notes
abstract
The 2014 i2b2 natural language processing shared task focused on identifying cardiovascular risk factors such as high blood pressure, high cholesterol levels, obesity and smoking status among other factors found in health records of diabetic patients. In addition, the task involved detecting medications, and time information associated with the extracted data. This paper presents the development and evaluation of a natural language processing (NLP) application conceived for this i2b2 shared task. For increased efficiency, the application main components were adapted from two existing NLP tools implemented in the Apache UIMA framework: Textractor (for dictionary-based lookup) and cTAKES (for preprocessing and smoking status detection). The application achieved a final (micro-averaged) F1-measure of 87.5% on the final evaluation test set. Our attempt was mostly based on existing tools adapted with minimal changes and allowed for satisfying performance with limited development efforts.
Abdulrahman Khalifa, Stéphane M. Meystre
J. Biomed. Informatics2
2014 Medication Prescription Status Classification in Clinical Narrative Documents
Jennifer H. Garvin, Julia Heavirland, Jenifer Williams, Stéphane M. Meystre
AMIA5
2014 Congestive Heart Failure Information Extraction Framework (CHIEF) Evaluation
Stéphane M. Meystre, Andrew Redd, Jennifer H. Garvin
AMIA1
2014 Effect of Pre-annotation on Annotation Time
Andrew Redd, Stéphane M. Meystre, Julia Heavirland, Allison Weaver, Jenifer Williams, Jennifer H. Garvin
AMIA3
2014 A System Usability Study Assessing a Machine-Assisted Interactive Interface to Support Annotation of Protected Health Information in Clinical Texts
Brett R. South, Danielle L. Mowery, Chris Leng, Stéphane M. Meystre, Wendy W. Chapman
AMIA4
2014 Text de-identification for privacy protection: A study of its impact on clinical text information content
Stéphane M. Meystre, Óscar Ferrández, F. Jeffrey Friedlin, Brett R. South, Shuying Shen, Matthew H. Samore
J. Biomed. Informatics1
2014 Evaluating the effects of machine pre-annotation and an interactive annotation interface on manual de-identification of clinical text
abstract
The Health Insurance Portability and Accountability Act (HIPAA) Safe Harbor method requires removal of 18 types of protected health information (PHI) from clinical documents to be considered "de-identified" prior to use for research purposes. Human review of PHI elements from a large corpus of clinical documents can be tedious and error-prone. Indeed, multiple annotators may be required to consistently redact information that represents each PHI class. Automated de-identification has the potential to improve annotation quality and reduce annotation time. For instance, using machine-assisted annotation by combining de-identification system outputs used as pre-annotations and an interactive annotation interface to provide annotators with PHI annotations for "curation" rather than manual annotation from "scratch" on raw clinical documents. In order to assess whether machine-assisted annotation improves the reliability and accuracy of the reference standard quality and reduces annotation effort, we conducted an annotation experiment. In this annotation study, we assessed the generalizability of the VA Consortium for Healthcare Informatics Research (CHIR) annotation schema and guidelines applied to a corpus of publicly available clinical documents called MTSamples. Specifically, our goals were to (1) characterize a heterogeneous corpus of clinical documents manually annotated for risk-ranked PHI and other annotation types (clinical eponyms and person relations), (2) evaluate how well annotators apply the CHIR schema to the heterogeneous corpus, (3) compare whether machine-assisted annotation (experiment) improves annotation quality and reduces annotation time compared to manual annotation (control), and (4) assess the change in quality of reference standard coverage with each added annotator's annotations.
Brett R. South, Danielle L. Mowery, Ying Suo, Jianwei Leng, Óscar Ferrández, Stéphane M. Meystre, Wendy W. Chapman
J. Biomed. Informatics6
2013 Automated Concept and Relationship Extraction for Ontology Development
Kristina Doing-Harris, Narong Boonsirisumpun, Kristi Potter, Yarden Livnat, Stéphane M. Meystre
AMIA5
2013 Semi-automated Ontology Development System for Medically Unexplained Syndromes in the U.S. Veterans Population
Stéphane M. Meystre, Kristina Doing-Harris, Narong Boonsirisumpun, Yarden Livnat, Kristi Potter
AMIA1
2013 Pediatric Acute Appendicitis Treatment Devices Automatic Extraction from Diagnostic Imaging Reports in a Multi-Institutional Clinical Repository
Stéphane M. Meystre, Ramkiran Gouripeddi, Abhisek Trivedi, Shawn Rangel
AMIA1
2013 BoB, a best-of-breed automated text de-identification system for VHA clinical documents
abstract
OBJECTIVE: De-identification allows faster and more collaborative clinical research while protecting patient confidentiality. Clinical narrative de-identification is a tedious process that can be alleviated by automated natural language processing methods. The goal of this research is the development of an automated text de-identification system for Veterans Health Administration (VHA) clinical documents. MATERIALS AND METHODS: We devised a novel stepwise hybrid approach designed to improve the current strategies used for text de-identification. The proposed system is based on a previous study on the best de-identification methods for VHA documents. This best-of-breed automated clinical text de-identification system (aka BoB) tackles the problem as two separate tasks: (1) maximize patient confidentiality by redacting as much protected health information (PHI) as possible; and (2) leave de-identified documents in a usable state preserving as much clinical information as possible. RESULTS: We evaluated BoB with a manually annotated corpus of a variety of VHA clinical notes, as well as with the 2006 i2b2 de-identification challenge corpus. We present evaluations at the instance- and token-level, with detailed results for BoB's main components. Moreover, an existing text de-identification system was also included in our evaluation. DISCUSSION: BoB's design efficiently takes advantage of the methods implemented in its pipeline, resulting in high sensitivity values (especially for sensitive PHI categories) and a limited number of false positives. CONCLUSIONS: Our system successfully addressed VHA clinical document de-identification, and its hybrid stepwise design demonstrates robustness and efficiency, prioritizing patient confidentiality while leaving most clinical information intact.
Óscar Ferrández, Brett R. South, Shuying Shen, F. Jeffrey Friedlin, Matthew H. Samore, Stéphane M. Meystre
J. Am. Medical Informatics Assoc.6
2012 Domain and Application Ontologies for Medically Unexplained Syndromes
Kristina Doing-Harris, Stéphane M. Meystre, Matthew H. Samore, Werner Ceusters
AMIA2
2012 Generalizability and Comparison of Automatic Clinical Text De-Identification Methods and Resources
Óscar Ferrández, Brett R. South, Shuying Shen, F. Jeffrey Friedlin, Matthew H. Samore, Stéphane M. Meystre
AMIA6
2012 Determining Section Types to Capture Key Clinical Data for Automation of Quality Measurement for Inpatients with Chronic Heart Failure
Jennifer H. Garvin, Julia Heavirland, Allison Weaver, Bruce E. Bray, Daniel Bolton, Carol Hope, Andrew Redd, Sarah A. Maulden, Stéphane M. Meystre
AMIA10
2012 A Survey of VHA Privacy Officers for the External Use of Automatically De-Identified Clinical Documents
Neil Nokes, Stéphane M. Meystre, Brett R. South, Jeffrey Scehnet, Shuying Shen, Óscar Ferrández, F. Jeffrey Friedlin, Matthew Maw, Matthew H. Samore
AMIA2
2012 On the Road Towards Developing a Publicly Available Corpus of De-identified Clinical Texts
Brett R. South, Danielle L. Mowery, Óscar Ferrández, Shuying Shen, Ying Suo, Annie Chen, Stéphane M. Meystre, Wendy W. Chapman
AMIA9
2012 Common data model for natural language processing based on two existing standard information models: CDA+GrAF
Stéphane M. Meystre, Chai Young Jung, Raphaël D. Chevrier
J. Biomed. Informatics1
2010 Textractor: a hybrid system for medications and reason for their prescription extraction from clinical text documents
abstract
UNLABELLED: OBJECTIVE To describe a new medication information extraction system-Textractor-developed for the 'i2b2 medication extraction challenge'. The development, functionalities, and official evaluation of the system are detailed. DESIGN: Textractor is based on the Apache Unstructured Information Management Architecture (UMIA) framework, and uses methods that are a hybrid between machine learning and pattern matching. Two modules in the system are based on machine learning algorithms, while other modules use regular expressions, rules, and dictionaries, and one module embeds MetaMap Transfer. MEASUREMENTS: The official evaluation was based on a reference standard of 251 discharge summaries annotated by all teams participating in the challenge. The metrics used were recall, precision, and the F(1)-measure. They were calculated with exact and inexact matches, and were averaged at the level of systems and documents. RESULTS: The reference metric for this challenge, the system-level overall F(1)-measure, reached about 77% for exact matches, with a recall of 72% and a precision of 83%. Performance was the best with route information (F(1)-measure about 86%), and was good for dosage and frequency information, with F(1)-measures of about 82-85%. Results were not as good for durations, with F(1)-measures of 36-39%, and for reasons, with F(1)-measures of 24-27%. CONCLUSION: The official evaluation of Textractor for the i2b2 medication extraction challenge demonstrated satisfactory performance. This system was among the 10 best performing systems in this challenge.
Stéphane M. Meystre, Julien Thibault, Shuying Shen, John F. Hurdle, Brett R. South
J. Am. Medical Informatics Assoc.1
2009 Detecting Intuitive Mentions of Diseases in Narrative Clinical Text
Stéphane M. Meystre
AIME1
2009 Inductive Creation of an Annotation Schema and a Reference Standard for De-identification of VA Electronic Clinical Notes
Jeanmarie Mayer, Shuying Shen, Brett R. South, Stéphane M. Meystre, F. Jeffrey Friedlin, William R. Ray, Matthew H. Samore
AMIA4
2009 A Clinical Use Case to Evaluate the i2b2 Hive: Predicting Asthma Exacerbations
Stéphane M. Meystre, Vikrant G. Deshmukh, Joyce A. Mitchell
AMIA1
2006 Improving the Sensitivity of the Problem List in an Intensive Care Unit by Using Natural Language Processing
Stéphane M. Meystre, Peter J. Haug
AMIA1
2006 Natural language processing to extract medical problems from electronic clinical documents: Performance evaluation
Stéphane M. Meystre, Peter J. Haug
J. Biomed. Informatics1
2005 Comparing Natural Language Processing Tools to Extract Medical Problems from Narrative Text
Stéphane M. Meystre, Peter J. Haug
AMIA1
2003 Medical Problem and Document Model for Natural Language Understanding
Stéphane M. Meystre, Peter J. Haug
AMIA1