VLDB 2026 Research / reviewers in the wild / expert
Christopher G. Chute
dblp:13/5552
· DBLP profile ↗
156ranked-venue papers
15as first author
16since 2021 · last 2026
0000-0001-5437-2545ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 150 · 15 first-author · 16 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Development of a robust corpus for automated evaluation of online health information in Chinese using the DISCERN scaleabstractOBJECTIVE: To develop the first comprehensive, standardized annotated corpus of Chinese online health information (OHI) using the full 16-item DISCERN instrument and to establish a reliable annotation process that supports automated quality assessment. MATERIALS AND METHODS: We assembled 510 web-sourced articles on breast cancer, arthritis, and depression. All the articles were independently annotated by three trained raters using the DISCERN scale. Annotation followed a four-step workflow: data collection and preprocessing, rater training, iterative annotation, and quality control. Raters calibrated through consensus sessions and calibration articles. The Dawid-Skene model aggregated individual annotations into final consensus scores. Original five-point ratings were retained and binarized (scores 1-3 as low quality, 4-5 as high quality) to enable both fine-grained and coarse evaluation for machine learning. RESULTS: Initial annotation of a 60-article pilot produced low agreement (mean Krippendorff's α ≈ 0.022) due to subjective variability. Successive calibration exercises improved agreement markedly, culminating in a corpus-wide Krippendorff's α of 0.834. Consensus scores correlated strongly with individual rater scores, confirming annotation robustness. The dual-scale design yielded a relatively balanced distribution of labels across topics, with roughly equal representation of low- and high-quality articles, and preserved granularity for detailed DISCERN analysis. DISCUSSION: Our iterative calibration approach and consensus modeling effectively addressed the subjective ambiguity inherent in quality assessment. The binary and five-class labeling strategies facilitate flexible downstream applications, allowing automated systems to perform both broad filtering and nuanced quality differentiation. The high inter-rater reliability demonstrates that rigorous training and consensus methods can overcome domain-specific annotation challenges. CONCLUSION: The resulting Chinese OHI corpus, annotated via a standardized DISCERN framework and refined through iterative calibration, provides a robust benchmark for training and evaluating machine learning models. This resource lays the foundation for scalable, reliable automated quality assessment of OHI in Chinese public health settings. Ting E, Xingxi Li, Junhao Ma, Qichuan Fang, Shanli Chen, Jianbo Lei, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 8 |
| 2025 | National COVID Cohort Collaborative data enhancements: a path for expanding common data modelsabstractOBJECTIVE: To support long COVID research in National COVID Cohort Collaborative (N3C), the N3C Phenotype and Data Acquisition team created data designs to aid contributing sites in enhancing their data. Enhancements include long COVID specialty clinic indicator; Admission, Discharge, and Transfer transactions; patient-level social determinants of health; and in-hospital use of oxygen supplementation. MATERIALS AND METHODS: For each enhancement, we defined the scope and wrote guidance on how to prepare and populate the data in a standardized way. RESULTS: As of June 2024, 29 sites have added at least one data enhancement to their N3C pipeline. DISCUSSION: The use of common data models is critical to the success of N3C; however, these data models cannot account for all needs. Project-driven data enhancement is required. This should be done in a standardized way in alignment with common data model specifications. Our approach offers a useful pathway for enhancing data to improve fit for purpose. CONCLUSION: In this initiative, we rapidly produced project-specific data modeling guidance and documentation in support of long COVID research while maintaining a commitment to terminology standards and harmonized data. Kellie M. Walters, Marshall Clark, Sofia Dard, Stephanie S. Hong, Elizabeth Kelly, Kristin Kostka, Adam M. Lee, Robert T. Miller, Michele Morris, Matvey Palchuk, Emily R. Pfaff, Adam B. Wilcox, Alexis Graves, Alfred Anzalone, Amin Manna, Amit Saha, Amy Olex, Andrea Zhou, Andrew E. Williams, Andrew Southerland, Andrew T. Girvin, Anita Walden, Anjali A Sharathkumar, Benjamin R. C. Amor, Benjamin Bates, Brian Hendricks, Caleb Alexander, Carolyn T. Bramante, Cavin Ward-Caviness, Charisse R. Madlock-Brown, Christine Suver, Christopher G. Chute, Christopher Dillon, Chunlei Wu, Clare Schmitt, Cliff Takemoto, Dan Housman, Davera Gabriel, David Eichmann, Diego Mazzotti, Don Brown, Eilis A. Boudreau, Elaine L. Hill, Elizabeth Zampino, Emily Carlson Marti, Evan French, Farrukh M. Koraishy, Federico Mariona, Fred W. Prior, George Sokos, Greg Martin, Harold P. Lehmann, Heidi Spratt, Hemalkumar Mehta, Hythem Sidky, J. W. Awori Hayanga, Jami Pincavitch, Jaylyn Clark, Jeremy Richard Harper, Jessica Islam, Jin Ge, Joel Gagnier, Joel H. Saltz, Johanna Loomba, John Buse, Jomol P. Mathew, Joni L. Rutter, Julie A. McMurry, Justin Guinney, Justin Starren, Karen Crowley, Katie Rebecca Bradwell, Ken Wilkins, Kenneth R. Gersing, Kenrick Dwain Cato, Kimberly Murray, Lavance Northington, Lee Allan Pyles, Leonie Misquitta, Lesley Cottrell, Lili M. Portilla, Mariam Deacy, Mark M. Bissell, Mary Emmett, Mary Morrison Saltz, Melissa A. Haendel, Meredith C. B. Adams, Meredith Temple-O'Connor, Michael G. Kurilla, Nabeel Qureshi, Nasia Safdar, Nicole Garbarini, Noha Sharafeldin, Ofer Sadan, Patricia A. Francis, Penny Wung Burgoon, Peter N. Robinson, Philip R. O. Payne, Rafael Fuentes, Randeep Jawa, Rebecca Erwin-Cohen, Rena Patel, Richard A. Moffitt, Richard L. Zhu, Rishi Kamaleswaran, Robert Hurley, Saiju Pyarajan, Samuel G. Michael, Samuel Bozzette, Sandeep Mallipattu, Satyanarayana Vedula, Scott Chapman, Shawn T. O'Neil, Soko Setoguchi, Tellen D. Bennett, Tiffany Callahan, Umit Topaloglu, Usman Sheikh, Valery Gordon, Vignesh Subbian, Warren A. Kibbe, Wenndy Hernandez, Will Beasley, Will Cooper, William Hillegass, Xiaohan Tanner Zhang |
J. Am. Medical Informatics Assoc. | 33 |
| 2023 | Characterizing variability of electronic health record-driven phenotype definitionsabstractOBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic. Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 15 |
| 2023 | Clinical encounter heterogeneity and methods for resolving in networked EHR data: a study from N3C and RECOVER programsabstractOBJECTIVE: Clinical encounter data are heterogeneous and vary greatly from institution to institution. These problems of variance affect interpretability and usability of clinical encounter data for analysis. These problems are magnified when multisite electronic health record (EHR) data are networked together. This article presents a novel, generalizable method for resolving encounter heterogeneity for analysis by combining related atomic encounters into composite "macrovisits." MATERIALS AND METHODS: Encounters were composed of data from 75 partner sites harmonized to a common data model as part of the NIH Researching COVID to Enhance Recovery Initiative, a project of the National Covid Cohort Collaborative. Summary statistics were computed for overall and site-level data to assess issues and identify modifications. Two algorithms were developed to refine atomic encounters into cleaner, analyzable longitudinal clinical visits. RESULTS: Atomic inpatient encounters data were found to be widely disparate between sites in terms of length-of-stay (LOS) and numbers of OMOP CDM measurements per encounter. After aggregating encounters to macrovisits, LOS and measurement variance decreased. A subsequent algorithm to identify hospitalized macrovisits further reduced data variability. DISCUSSION: Encounters are a complex and heterogeneous component of EHR data and native data issues are not addressed by existing methods. These types of complex and poorly studied issues contribute to the difficulty of deriving value from EHR data, and these types of foundational, large-scale explorations, and developments are necessary to realize the full potential of modern real-world data. CONCLUSION: This article presents method developments to manipulate and resolve EHR encounter data issues in a generalizable way as a foundation for future research and analysis. Peter Leese, Adit Anand, Andrew T. Girvin, Amin Manna, Saaya Patel, Yun Jae Yoo, Rachel Wong, Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, Janos G. Hajagos, Emily R. Pfaff, Richard A. Moffitt |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)abstractDespite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts. Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 22 |
| 2023 | De-black-boxing health AI: demonstrating reproducible machine learning computable phenotypes using the N3C-RECOVER Long COVID model in the All of Us data repositoryabstractMachine learning (ML)-driven computable phenotypes are among the most challenging to share and reproduce. Despite this difficulty, the urgent public health considerations around Long COVID make it especially important to ensure the rigor and reproducibility of Long COVID phenotyping algorithms such that they can be made available to a broad audience of researchers. As part of the NIH Researching COVID to Enhance Recovery (RECOVER) Initiative, researchers with the National COVID Cohort Collaborative (N3C) devised and trained an ML-based phenotype to identify patients highly probable to have Long COVID. Supported by RECOVER, N3C and NIH's All of Us study partnered to reproduce the output of N3C's trained model in the All of Us data enclave, demonstrating model extensibility in multiple environments. This case study in ML-based phenotype reuse illustrates how open-source software best practices and cross-site collaboration can de-black-box phenotyping algorithms, prevent unnecessary rework, and promote open science in informatics. Emily R. Pfaff, Andrew T. Girvin, Miles Crosskey, Srushti Gangireddy, Hiral Master, Wei-Qi Wei, Vern Eric Kerchberger, Mark G. Weiner, Paul A. Harris, Melissa A. Basford, Chris Lunt, Christopher G. Chute, Richard A. Moffitt, Melissa A. Haendel |
J. Am. Medical Informatics Assoc. | 12 |
| 2022 | Making a Case for Adoption of ICD-11 Morbidity Reporting in the U.S.: Multiple Perspectives
Christopher G. Chute, Susan H. Fenton, Mary H. Stanfill, Kathy L. Giannangelo |
AMIA | 1 |
| 2022 | Modeling a Cancer Symptom Control Domain Using HL7 FHIR: Applicability of the Minimal Common Oncology Data Elements (mCODE)
Nan Huo, Yue Yu 0012, Nansu Zong, Andrea Cheville, Claude J. Nanjo, Eric Prud'hommeaux, Deirdre Pachman, Guohui Xiao 0001, Emily R. Pfaff, Christopher G. Chute, Guoqian Jiang, Kathryn J. Ruddy |
AMIA | 10 |
| 2022 | A Comparative Study on the Capability of Real-World Antineoplastic Drug Data Collection by CanMED, ATC and HemOnc
Yue Yu 0012, Kathryn J. Ruddy, Nan Huo, Nansu Zong, Deirdre Pachman, Christopher G. Chute, Emily R. Pfaff, Andrea Cheville, Guoqian Jiang |
AMIA | 6 |
| 2022 | Analyzing historical diagnosis code data from NIH N3C and RECOVER Programs using deep learning to determine risk factors for Long CovidabstractPost-acute sequelae of SARS-CoV-2 infection (PASC) or Long COVID is an emerging medical condition that has been observed in several patients with a positive diagnosis for COVID-19. Historical Electronic Health Records (EHR) like diagnosis codes, lab results and clinical notes have been analyzed using deep learning and have been used to predict future clinical events. In this paper, we propose an interpretable deep learning approach to analyze historical diagnosis code data from the National COVID Cohort Collective (N3C)1to find the risk factors contributing to developing Long COVID. Using our deep learning approach, we are able to predict if a patient is suffering from Long COVID from a temporally ordered list of diagnosis codes up to 45 days post the first COVID positive test or diagnosis for each patient, with an accuracy of 70.48%. We are then able to examine the trained model using Gradient-weighted Class Activation Mapping (GradCAM) to give each input diagnoses a score. The highest scored diagnosis were deemed to be the most important for making the correct prediction for a patient. We also propose a way to summarize these top diagnoses for each patient in our cohort and look at their temporal trends to determine which codes contribute towards a positive Long COVID diagnosis. Saurav Sengupta, Johanna Loomba, Suchetha Sharma, Donald E. Brown, Lorna E. Thorpe, Melissa A. Haendel, Christopher G. Chute, Stephanie S. Hong |
BIBM | 7 |
| 2022 | Harmonizing units and values of quantitative data elements in a very large nationally pooled electronic health record (EHR) datasetabstractOBJECTIVE: The goals of this study were to harmonize data from electronic health records (EHRs) into common units, and impute units that were missing. MATERIALS AND METHODS: The National COVID Cohort Collaborative (N3C) table of laboratory measurement data-over 3.1 billion patient records and over 19 000 unique measurement concepts in the Observational Medical Outcomes Partnership (OMOP) common-data-model format from 55 data partners. We grouped ontologically similar OMOP concepts together for 52 variables relevant to COVID-19 research, and developed a unit-harmonization pipeline comprised of (1) selecting a canonical unit for each measurement variable, (2) arriving at a formula for conversion, (3) obtaining clinical review of each formula, (4) applying the formula to convert data values in each unit into the target canonical unit, and (5) removing any harmonized value that fell outside of accepted value ranges for the variable. For data with missing units for all the results within a lab test for a data partner, we compared values with pooled values of all data partners, using the Kolmogorov-Smirnov test. RESULTS: Of the concepts without missing values, we harmonized 88.1% of the values, and imputed units for 78.2% of records where units were absent (41% of contributors' records lacked units). DISCUSSION: The harmonization and inference methods developed herein can serve as a resource for initiatives aiming to extract insight from heterogeneous EHR collections. Unique properties of centralized data are harnessed to enable unit inference. CONCLUSION: The pipeline we developed for the pooled N3C data enables use of measurements that would otherwise be unavailable for analysis. Katie R. Bradwell, Jacob T. Wooldridge, Benjamin R. C. Amor, Tellen D. Bennett, Adit Anand, Carolyn Bremer, Yun Jae Yoo, Zhenglong Qian, Steven G. Johnson, Emily R. Pfaff, Andrew T. Girvin, Amin Manna, Emily Niehaus, Stephanie S. Hong, Xiaohan Tanner Zhang, Richard L. Zhu, Mark Bissell, Nabeel Qureshi, Joel H. Saltz, Melissa A. Haendel, Christopher G. Chute, Harold P. Lehmann, Richard A. Moffitt |
J. Am. Medical Informatics Assoc. | 21 |
| 2022 | Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborativeabstractOBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require. Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai |
J. Am. Medical Informatics Assoc. | 22 |
| 2022 | Demonstrating an approach for evaluating synthetic geospatial and temporal epidemiologic data utility: results from analyzing >1.8 million SARS-CoV-2 tests in the United States National COVID Cohort Collaborative (N3C)abstractOBJECTIVE: This study sought to evaluate whether synthetic data derived from a national coronavirus disease 2019 (COVID-19) dataset could be used for geospatial and temporal epidemic analyses. MATERIALS AND METHODS: Using an original dataset (n = 1 854 968 severe acute respiratory syndrome coronavirus 2 tests) and its synthetic derivative, we compared key indicators of COVID-19 community spread through analysis of aggregate and zip code-level epidemic curves, patient characteristics and outcomes, distribution of tests by zip code, and indicator counts stratified by month and zip code. Similarity between the data was statistically and qualitatively evaluated. RESULTS: In general, synthetic data closely matched original data for epidemic curves, patient characteristics, and outcomes. Synthetic data suppressed labels of zip codes with few total tests (mean = 2.9 ± 2.4; max = 16 tests; 66% reduction of unique zip codes). Epidemic curves and monthly indicator counts were similar between synthetic and original data in a random sample of the most tested (top 1%; n = 171) and for all unsuppressed zip codes (n = 5819), respectively. In small sample sizes, synthetic data utility was notably decreased. DISCUSSION: Analyses on the population-level and of densely tested zip codes (which contained most of the data) were similar between original and synthetically derived datasets. Analyses of sparsely tested populations were less similar and had more data suppression. CONCLUSION: In general, synthetic data were successfully used to analyze geospatial and temporal trends. Analyses using small sample sizes or populations were limited, in part due to purposeful data label suppression-an attribute disclosure countermeasure. Users should consider data fitness for use in these cases. Jason A. Thomas, Randi E. Foraker, Noa Zamstein, Jon D. Morrow, Philip R. O. Payne, Adam B. Wilcox, Melissa A. Haendel, Christopher G. Chute, Kenneth R. Gersing, Anita Walden, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Justin Starren, Christine Suver, Chunlei Wu, Davera Gabriel, Stephanie S. Hong, Kristin Kostka, Harold P. Lehmann, Richard A. Moffitt, Michele Morris, Matvey Palchuk, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Mark M. Bissell, Marshall Clark, Andrew T. Girvin, Adam M. Lee, Robert T. Miller, Kellie M. Walters, Yooree Chae, Connor Cook, Alexandra Dest, Racquel R. Dietz, Thomas Dillon, Patricia A. Francis, Rafael Fuentes, Alexis Graves, Andrew J. Neumann, Shawn T. O'Neil, Usman Sheikh, Andréa M. Volz, Elizabeth Zampino, Christopher P. Austin, Samuel Bozzette, Mariam Deacy, Nicole Garbarini, Michael G. Kurilla, Samuel G. Michael, Joni L. Rutter, Meredith Temple-O'Connor, Katie Rebecca Bradwell, Amin Manna, Nabeel Qureshi, Mary Morrison Saltz, Julie A. McMurry, Carolyn T. Bramante, Jeremy Richard Harper, Wenndy Hernandez, Farrukh M. Koraishy, Federico Mariona, Saidulu Mattapally, Amit Saha, Satyanarayana Vedula, Yujuan Fu, Nisha Mathews, Ofer Mendelevitch |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | FHIR-Ontop-OMOP: Building clinical knowledge graphs in FHIR RDF with the OMOP Common data ModelabstractBACKGROUND: Knowledge graphs (KGs) play a key role to enable explainable artificial intelligence (AI) applications in healthcare. Constructing clinical knowledge graphs (CKGs) against heterogeneous electronic health records (EHRs) has been desired by the research and healthcare AI communities. From the standardization perspective, community-based standards such as the Fast Healthcare Interoperability Resources (FHIR) and the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) are increasingly used to represent and standardize EHR data for clinical data analytics, however, the potential of such a standard on building CKG has not been well investigated. OBJECTIVE: To develop and evaluate methods and tools that expose the OMOP CDM-based clinical data repositories into virtual clinical KGs that are compliant with FHIR Resource Description Framework (RDF) specification. METHODS: We developed a system called FHIR-Ontop-OMOP to generate virtual clinical KGs from the OMOP relational databases. We leveraged an OMOP CDM-based Medical Information Mart for Intensive Care (MIMIC-III) data repository to evaluate the FHIR-Ontop-OMOP system in terms of the faithfulness of data transformation and the conformance of the generated CKGs to the FHIR RDF specification. RESULTS: A beta version of the system has been released. A total of more than 100 data element mappings from 11 OMOP CDM clinical data, health system and vocabulary tables were implemented in the system, covering 11 FHIR resources. The generated virtual CKG from MIMIC-III contains 46,520 instances of FHIR Patient, 716,595 instances of Condition, 1,063,525 instances of Procedure, 24,934,751 instances of MedicationStatement, 365,181,104 instances of Observations, and 4,779,672 instances of CodeableConcept. Patient counts identified by five pairs of SQL (over the MIMIC database) and SPARQL (over the virtual CKG) queries were identical, ensuring the faithfulness of the data transformation. Generated CKG in RDF triples for 100 patients were fully conformant with the FHIR RDF specification. CONCLUSION: The FHIR-Ontop-OMOP system can expose OMOP database as a FHIR-compliant RDF graph. It provides a meaningful use case demonstrating the potentials that can be enabled by the interoperability between FHIR and OMOP CDM. Generated clinical KGs in FHIR RDF provide a semantic foundation to enable explainable AI applications in healthcare. Guohui Xiao 0001, Emily R. Pfaff, Eric Prud'hommeaux, David Booth, Deepak K. Sharma, Nan Huo, Yue Yu 0012, Nansu Zong, Kathryn J. Ruddy, Christopher G. Chute, Guoqian Jiang |
J. Biomed. Informatics | 10 |
| 2022 | Developing an ETL tool for converting the PCORnet CDM into the OMOP CDM to facilitate the COVID-19 data integration
Yue Yu 0012, Nansu Zong, Andrew Wen, Sijia Liu 0002, Daniel J. Stone, David Knaack, Alanna M. Chamberlain, Emily R. Pfaff, Davera Gabriel, Christopher G. Chute, Nilay Shah, Guoqian Jiang |
J. Biomed. Informatics | 10 |
| 2021 | The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deploymentabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19. Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Toward a Harmonized WHO Family of International Classifications Content Model
Samson W. Tu, Csongor Nyulas, Tania Tudorache, Mark A. Musen, Andrea Martinuzzi, Coen H. van Gool, Vincenzo Della Mea, Christopher G. Chute, Lucilla Frattura, Nicholas R. Hardiker, Huib ten Napel, Richard Madden, Ann-Helene Almborg, Jeewani Anupama Ginige, Catherine Sykes, Can Çelik, Robert Jakob |
AMIA | 8 |
| 2019 | The 2018 fellow cohort of the American College of Medical InformaticsabstractFounded in 1984, the American College of Medical Informatics (ACMI) is a college of elected Fellows from the United States and abroad who have made significant and sustained contributions to the field of biomedical informatics. On November 4, 2018, the 2018 Cohort of Fellows was introduced to the College and attendees at the American Medical Informatics Association (AMIA) Annual Symposium as well as to the public through Tweets. This article includes the introduction for each Fellow in the 2018 Cohort, which was read by Christopher G. Chute (ACMI President), Suzanne Bakken (ACMI Past President), or William M. Tierney (ACMI President-Elect); a somewhat tongue-in-cheek Tweet created by Gretchen Purcell Jackson or James J. Cimino; and a link to their Journal of theAmerican Medical Informatics Association (JAMIA) and JAMIA Open publications. This is followed by the traditional closing remarks of welcome into ACMI. Gregory L. Alexander,PhD, RN, FAAN, FACMI Potter-Brinton Endowed Professor, Sinclair School of Nursing and Department of Health Management and Informatics, University of Missouri Christopher G. Chute, Suzanne Bakken, William M. Tierney, Gretchen Purcell Jackson, James J. Cimino |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Sex, obesity, diabetes, and exposure to particulate matter among patients with severe asthma: Scientific insights from a comparative analysis of open clinical data sources during a five-day hackathonabstractThis special communication describes activities, products, and lessons learned from a recent hackathon that was funded by the National Center for Advancing Translational Sciences via the Biomedical Data Translator program ('Translator'). Specifically, Translator team members self-organized and worked together to conceptualize and execute, over a five-day period, a multi-institutional clinical research study that aimed to examine, using open clinical data sources, relationships between sex, obesity, diabetes, and exposure to airborne fine particulate matter among patients with severe asthma. The goal was to develop a proof of concept that this new model of collaboration and data sharing could effectively produce meaningful scientific results and generate new scientific hypotheses. Three Translator Clinical Knowledge Sources, each of which provides open access (via Application Programming Interfaces) to data derived from the electronic health record systems of major academic institutions, served as the source of study data. Jupyter Python notebooks, shared in GitHub repositories, were used to call the knowledge sources and analyze and integrate the results. The results replicated established or suspected relationships between sex, obesity, diabetes, exposure to airborne fine particulate matter, and severe asthma. In addition, the results demonstrated specific differences across the three Translator Clinical Knowledge Sources, suggesting cohort- and/or environment-specific factors related to the services themselves or the catchment area from which each service derives patient data. Collectively, this special communication demonstrates the power and utility of intense, team-oriented hackathons and offers general technical, organizational, and scientific lessons learned. Karamarie Fecho, Stanley C. Ahalt, Saravanan Arunachalam, James Champion, Christopher G. Chute, Sarah Davis, Kenneth Gersing, Gwênlyn Glusman, Jennifer Hadlock, Jewel Lee, Emily R. Pfaff, Max Robinson, Eric Sid, Casey N. Ta, Hao Xu 0006, Richard L. Zhu, Qian Zhu 0003, David B. Peden |
J. Biomed. Informatics | 5 |
| 2018 | Characterizing Design Patterns of EHR-Driven Phenotype Extraction Algorithms
Yizhen Zhong, Luke V. Rasmussen, Jennifer A. Pacheco, Maureen E. Smith, Justin Starren, Wei-Qi Wei, Peter Speltz, Joshua C. Denny, Nephi Walton, George Hripcsak, Christopher G. Chute, Yuan Luo 0001 |
BIBM | 12 |
| 2018 | Empowering genomic medicine by establishing critical sequencing result data flows: the eMERGE exampleabstractThe eMERGE Network is establishing methods for electronic transmittal of patient genetic test results from laboratories to healthcare providers across organizational boundaries. We surveyed the capabilities and needs of different network participants, established a common transfer format, and implemented transfer mechanisms based on this format. The interfaces we created are examples of the connectivity that must be instantiated before electronic genetic and genomic clinical decision support can be effectively built at the point of care. This work serves as a case example for both standards bodies and other organizations working to build the infrastructure required to provide better electronic clinical decision support for clinicians. Samuel J. Aronson, Lawrence J. Babb, Darren C. Ames, Richard A. Gibbs, Eric Venner, John J. Connelly, Keith Marsolo, Chunhua Weng, Marc S. Williams, Andrea L. Hartzler, Wayne H. Liang, James D. Ralston, Emily Beth Devine, Shawn N. Murphy, Christopher G. Chute, Pedro J. Caraballo, Iftikhar J. Kullo, Robert R. Freimuth, Luke V. Rasmussen, Firas H. Wehbe, Josh F. Peterson, Jamie R. Robinson, Ken Wiley, Casey Overby Taylor |
J. Am. Medical Informatics Assoc. | 15 |
| 2017 | Value of Genetics-informed Drug Dosing Guidance in Pregnant Women: A Needs Assessment with Obstetric Healthcare Providers at Johns Hopkins
Casey Overby Taylor, Phillip Thompkins, Harold P. Lehmann, Christopher G. Chute, Jeanne Sheffield |
AMIA | 4 |
| 2017 | SMART-on-FHIR implemented over i2b2abstractWe have developed an interface to serve patient data from Informatics for Integrating Biology and the Bedside (i2b2) repositories in the Fast Healthcare Interoperability Resources (FHIR) format, referred to as a SMART-on-FHIR cell. The cell serves FHIR resources on a per-patient basis, and supports the "substitutable" modular third-party applications (SMART) OAuth2 specification for authorization of client applications. It is implemented as an i2b2 server plug-in, consisting of 6 modules: authentication, REST, i2b2-to-FHIR converter, resource enrichment, query engine, and cache. The source code is freely available as open source. We tested the cell by accessing resources from a test i2b2 installation, demonstrating that a SMART app can be launched from the cell that accesses patient data stored in i2b2. We successfully retrieved demographics, medications, labs, and diagnoses for test patients. The SMART-on-FHIR cell will enable i2b2 sites to provide simplified but secure data access in FHIR format, and will spur innovation and interoperability. Further, it transforms i2b2 into an apps platform. Kavishwar B. Wagholikar, Joshua C. Mandel, Jeffrey G. Klann, Nich Wattanasin, Michael Mendis, Christopher G. Chute, Kenneth D. Mandl, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 6 |
| 2016 | Clinical element models in the SHARPn consortiumabstractOBJECTIVE: The objective of the Strategic Health IT Advanced Research Project area four (SHARPn) was to develop open-source tools that could be used for the normalization of electronic health record (EHR) data for secondary use--specifically, for high throughput phenotyping. We describe the role of Intermountain Healthcare's Clinical Element Models ([CEMs] Intermountain Healthcare Health Services, Inc, Salt Lake City, Utah) as normalization "targets" within the project. MATERIALS AND METHODS: Intermountain's CEMs were either repurposed or created for the SHARPn project. A CEM describes "valid" structure and semantics for a particular kind of clinical data. CEMs are expressed in a computable syntax that can be compiled into implementation artifacts. The modeling team and SHARPn colleagues agilely gathered requirements and developed and refined models. RESULTS: Twenty-eight "statement" models (analogous to "classes") and numerous "component" CEMs and their associated terminology were repurposed or developed to satisfy SHARPn high throughput phenotyping requirements. Model (structural) mappings and terminology (semantic) mappings were also created. Source data instances were normalized to CEM-conformant data and stored in CEM instance databases. A model browser and request site were built to facilitate the development. DISCUSSION: The modeling efforts demonstrated the need to address context differences and granularity choices and highlighted the inevitability of iso-semantic models. The need for content expertise and "intelligent" content tooling was also underscored. We discuss scalability and sustainability expectations for a CEM-based approach and describe the place of CEMs relative to other current efforts. CONCLUSIONS: The SHARPn effort demonstrated the normalization and secondary use of EHR data. CEMs proved capable of capturing data originating from a variety of sources within the normalization pipeline and serving as suitable normalization targets. Thomas A. Oniki, Ning Zhuo, Calvin E. Beebe, Joseph F. Coyle, Craig G. Parker, Harold R. Solbrig, Kyle Marchant, Vinod Kaggal, Christopher G. Chute, Stanley M. Huff |
J. Am. Medical Informatics Assoc. | 10 |
| 2016 | Developing a data element repository to support EHR-driven phenotype algorithm authoring and execution
Guoqian Jiang, Richard C. Kiefer, Luke V. Rasmussen, Harold R. Solbrig, Huan Mo, Jennifer A. Pacheco, Jie Xu 0011, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
J. Biomed. Informatics | 11 |
| 2015 | Harmonization of Quality Data Model with HL7 FHIR to Support EHR-driven Phenotype Authoring and Execution: A Pilot Study
Guoqian Jiang, Harold R. Solbrig, Richard C. Kiefer, Luke V. Rasmussen, Huan Mo, Jennifer A. Pacheco, Enid N. H. Montague, Jie Xu 0011, Peter Speltz, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
AMIA | 12 |
| 2015 | Quality Assurance of Cancer Study Common Data Elements Using A Post-Coordination Approach
Guoqian Jiang, Harold R. Solbrig, Eric Prud'hommeaux, Cui Tao, Chunhua Weng, Christopher G. Chute |
AMIA | 6 |
| 2015 | Harmonization of ICD-11 and SNOMED CT - Not just mapping! Practical and Theoretical Lessons & Benefits to Users and Implementers
Alan L. Rector, James R. Campbell 0001, Bedirhan Üstün, Christopher G. Chute, Harold R. Solbrig |
AMIA | 4 |
| 2015 | Representing and Validating Cancer Study Metadata Standard Using RDF Shapes Expression Language
Harold R. Solbrig, Eric Prud'hommeaux, Christopher G. Chute, Guoqian Jiang |
AMIA | 3 |
| 2015 | Transformation of standardized clinical models based on OWL technologies: from CEM to OpenEHR archetypesabstractINTRODUCTION: The semantic interoperability of electronic healthcare records (EHRs) systems is a major challenge in the medical informatics area. International initiatives pursue the use of semantically interoperable clinical models, and ontologies have frequently been used in semantic interoperability efforts. The objective of this paper is to propose a generic, ontology-based, flexible approach for supporting the automatic transformation of clinical models, which is illustrated for the transformation of Clinical Element Models (CEMs) into openEHR archetypes. METHODS: Our transformation method exploits the fact that the information models of the most relevant EHR specifications are available in the Web Ontology Language (OWL). The transformation approach is based on defining mappings between those ontological structures. We propose a way in which CEM entities can be transformed into openEHR by using transformation templates and OWL as common representation formalism. The transformation architecture exploits the reasoning and inferencing capabilities of OWL technologies. RESULTS: We have devised a generic, flexible approach for the transformation of clinical models, implemented for the unidirectional transformation from CEM to openEHR, a series of reusable transformation templates, a proof-of-concept implementation, and a set of openEHR archetypes that validate the methodological approach. CONCLUSIONS: We have been able to transform CEM into archetypes in an automatic, flexible, reusable transformation approach that could be extended to other clinical model specifications. We exploit the potential of OWL technologies for supporting the transformation process. We believe that our approach could be useful for international efforts in the area of semantic interoperability of EHR systems. María Del Carmen Legaz-García, Marcos Menárguez Tortosa, Jesualdo Tomás Fernández-Breis, Christopher G. Chute, Cui Tao |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 18 |
| 2015 | Health information technology data standards get down to business: maturation within domains and the emergence of interoperabilityabstractHealth information technology (HIT) standards are not new. Arguably, they date to the canonical list for causes of death in the London Bills of Mortality of 1528, which was later formalized during the middle of the 19th century into what we now recognize as the International Classifications of Diseases (ICD). Beginning in the middle of the 20th century, HIT standards evolved beyond vital statistics and began to capture data related to clinical morbidity, thereby facilitating nascent decision support, outcomes research, evidence generation, and health care quality improvement initiatives. Alas, with that expansion came a proliferation of competing and overlapping standards, giving substance to the critical aphorism that “the only nice thing about standards is that there are so many to choose from.” The emergence of large-scale computerization throughout health care in the last half century has further accelerated this divergence of standards, creating a veritable cacophony of noninteroperable medical record content, data exchange formalisms, and data silos. This special issue on data standards was prompted by a palpable maturation among HIT standards in clinical practice and biomedical research in just the last decade. There has been remarkable cooperation among HIT standards development organizations, including the new agreement to harmonize and coordinate overlapping content in Systematized Nomenclature of Medicine—Clinical Terms (SNOMED CT) and Logical Observation Identifiers Names and Codes (LOINC), and the historic cooperation between the SNOMED CT and ICD developers to create ICD11 on the semantic foundation of SNOMED CT. In parallel, there have also been unprecedented consolidation and harmonization of orthogonal standards into an emerging suite of specifications for health and biomedical observations such as the ONC Meaningful Use and the NIH Common Data Element efforts within the United States. While far from comprehensive or fully coherent, the current state of HIT standards is at a turning point, where we appear to be making more effective progress and practical applications than most would have predicted from the bad old days of just a few decades ago. The goal of this special focus issue of the JAMIA is to provide a forum for the latest evaluation of HIT standards in contexts of Meaningful Use, biomedical research, and big data. JAMIA editors solicited this focus issue with the hope that HIT standards and recent consolidation initiatives had indeed adequately matured so that the informatics community would respond with successful demonstrations for how particular standards can and will effectively support biomedical and population health applications. We thus broadly solicited scholarly contributions that would address evaluation, application, consolidation, or domain extensions of biomedical data standards within the framework of biomedical informatics, and we explicitly requested authors to provide supporting data, rigorous evaluation, or evidence of relevant consensus to support their work. We were not disappointed and enjoyed the opportunity to review many rich submissions for this highly competitive volume. We present in this issue the best of these submissions, collectively covering wide spectra of domains, applications, and approaches. The selected articles represent the full continuum of molecular, clinical, organizational, and population data, address a range of objectives from evaluation to application, and illustrate multiple approaches to HIT standards and interoperability from domain-specific to broad integration. Exemplifying the increasing coordination between multiple standards is the characterization of genetic data, as standardized by the Human Gene Nomenclature Committee with LOINC laboratory reports for genetic data.1 This represents the reuse of existing standards for genomic specification within an established framework for health data exchange and messaging. In work that also binds the basic science world to clinical practice, this history and success of the Human Proteome Organization Proteomics Standards Initiative details the broadly-based collaborations that have led to an increasingly mature and practical specification of proteomic findings.2 Similarly, poison control centers have demonstrated how they can collaborate with emergency departments by adopting their reference model for health information exchange to work within the HL7 Consolidated Clinical Document Architecture standard.3 This is an elegant demonstration of a community with an important clinical data exchange requirement choosing to embrace and enable an emerging mainstream mechanism, as opposed to the traditional solution of building yet another syntax to implement their reference model. On a more abstract level is the description of how the venerable Clinical Element Models can be systematically transformed into the openEHR archetype representations by invoking shared Web Ontology Language (OWL) formalisms.4 This demonstrates how the Clinical Information Modeling Initiative, through fostering community consensus on modeling syntax and languages, can stimulate informatics work that can harvest legacy content while strengthening standards harmonization and coherency. Demonstrating that not all standards are final or even fully robust, methods for discovering errors in large-scale terminologies, specifically SNOMED CT, substantially enhance the scale and completeness of quality assurance work within complex data standards.5 Such work has consistently demonstrated room for improvement in these large data specifications, and these standards together with their user communities are by far better for it. Correspondingly, a systemic review of the HL7/LOINC Document Ontology Role Axis highlights many shortcomings in representing clinical-role behaviors that are encountered in real-world resources.6 Such work, again, can and will feed back to the standards developers, making these components more robust within suites of HIT data standards. The pattern of empirical discovery is also evident in the systematic “bottom-up” evaluation of nursing documents across many standards developing organizations (SDO) contributors.7 Specifically, in the generation of eMeasures, clear and specific refinements of many clinical standards are needed and will, no doubt, come about because of these careful, reality-based evaluations. Continuing within the nursing domain, investigators demonstrate how collaboration across organizations working with established standards enables detailed and informatics messaging around hospital-acquired pressure ulcers.8 These efforts can directly address ONC challenges for the development of mobile applications to improve clinical care around this all too frequent complication of chronic care. Expanding from the specific to the strategic, a consensus community developed national action plans for collecting comparable nursing data to support secondary use as well as clinical care.9 By and large, nursing data is not yet systematically integrated into electronic health records, impoverishing both the care process and outcomes research. This national action plan promises to advance the evaluation and practice of nursing by supporting the generation of comparable nursing data for quality reporting and translational research. Domain-specific adaptations of existing data models and semantics obviously can extend to many other areas of health care. An evaluation of an oncology-specific implementation guide of the consolidated clinical data architecture standard is described as a case example, which completed the full HL7 ballot process.10 The paper also describes its clinical implementation by two organizations, which validated the overall strategy of domain-specific adaptations of established HIT standards. In a related domain-specific HL7 effort, the maturation of data elements for emergency departments is described.11 This process also involved the use and mapping of many related data standards. Meanwhile, the critical evaluation of semantics and value set binding to clinical models in the domain of heart failure demonstrates a pathway for semantically accurate interoperability in complex clinical domains.12 They were able to simplify that complexity into a framework of semantic patterns, enabling coherent access to heterogeneous data resources. Finally, the harmonization of distributed, heterogeneous clinical data into a shared specification based on HL7’s Virtual Medical Record specification illustrates how data standards can help integrate existing data, even if those data were not collected with fully specified or shared standards specifications.13 This paper shows how meta-standards can facilitate the integration of patient information from heterogeneous sources, a fundamental and all too common requirement for clinical decision support systems in the United States. Taken together, the papers presented in this special focus issue inform us about the state of HIT standards and their evolution. Clearly, they highlight that far more work remains, but more pertinently the overarching message is that collaboration, coordination, and convergence is occurring and may even be considered as the default effort. This is in vast contradiction to an earlier era when virtually every biomedical data standards development organization felt obligated to publish a competing “me too” standard, if only so that the products of their competitors would not succeed. We have all grown beyond that. HIT standards are now widely recognized not as commercial ventures in and of themselves but as critical public resources that can stimulate innovation and support applications that impact provider behaviors and patient outcomes. Efforts today are clearly focused on effective communication, semantic consistency, and interoperability. We can be sanguine about the likely state we may find ourselves in within the next decade. We have arrived at an era of real progress. Albeit slow, the maturation of practical, coherent, and interoperable biomedical data standards is undeniable and bodes well for clinical data interoperability. Rachel L. Richesson, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | CSER and eMERGE: current and potential state of the display of genetic information in the electronic health recordabstractOBJECTIVE: Clinicians' ability to use and interpret genetic information depends upon how those data are displayed in electronic health records (EHRs). There is a critical need to develop systems to effectively display genetic information in EHRs and augment clinical decision support (CDS). MATERIALS AND METHODS: The National Institutes of Health (NIH)-sponsored Clinical Sequencing Exploratory Research and Electronic Medical Records & Genomics EHR Working Groups conducted a multiphase, iterative process involving working group discussions and 2 surveys in order to determine how genetic and genomic information are currently displayed in EHRs, envision optimal uses for different types of genetic or genomic information, and prioritize areas for EHR improvement. RESULTS: There is substantial heterogeneity in how genetic information enters and is documented in EHR systems. Most institutions indicated that genetic information was displayed in multiple locations in their EHRs. Among surveyed institutions, genetic information enters the EHR through multiple laboratory sources and through clinician notes. For laboratory-based data, the source laboratory was the main determinant of the location of genetic information in the EHR. The highest priority recommendation was to address the need to implement CDS mechanisms and content for decision support for medically actionable genetic information. CONCLUSION: Heterogeneity of genetic information flow and importance of source laboratory, rather than clinical content, as a determinant of information representation are major barriers to using genetic information optimally in patient care. Greater effort to develop interoperable systems to receive and consistently display genetic and/or genomic information and alert clinicians to genomic-dependent improvements to clinical care is recommended. Brian H. Shirts, Joseph S. Salama, Samuel J. Aronson, Wendy K. Chung, Stacy W. Gray, Lucia Hindorff, Gail P. Jarvik, Sharon E. Plon, Elena M. Stoffel, Peter Tarczy-Hornoch, Eliezer M. Van Allen, Karen E. Weck, Christopher G. Chute, Robert R. Freimuth, Robert Grundmeier, Andrea L. Hartzler, Rongling Li, Peggy L. Peissig, Josh F. Peterson, Luke V. Rasmussen, Justin Starren, Marc S. Williams, Casey Overby Taylor |
J. Am. Medical Informatics Assoc. | 13 |
| 2014 | Evaluation of RxNorm for Medication Clinical Decision Support
Robert R. Freimuth, Kelly Wix, Qian Zhu 0003, Mark Siska, Christopher G. Chute |
AMIA | 5 |
| 2014 | Developing a Section Labeler for Clinical Documents
Peter J. Haug, Xinzi Wu, Jeffrey P. Ferraro, Guergana K. Savova, Stanley M. Huff, Christopher G. Chute |
AMIA | 6 |
| 2014 | Lexical Term Standardization of ICD-11 Using Semantic Web Technologies
Guoqian Jiang, Harold R. Solbrig, Bedirhan Üstün, Christopher G. Chute |
AMIA | 4 |
| 2014 | An Automated Approach for Ranking Journals to Help in Clinician Decision Support
Siddhartha Jonnalagadda, Soheil Moosavinasab, Chinmoy Nath, Dingcheng Li, Christopher G. Chute |
AMIA | 5 |
| 2014 | A Template for Authoring and Adapting Genomic Medicine Content in the eMERGE Infobutton Project
Casey Overby Taylor, Luke V. Rasmussen, Andrea L. Hartzler, John J. Connolly, Josh F. Peterson, RoseMary Hedberg, Robert R. Freimuth, Brian H. Shirts, Joshua C. Denny, Eric B. Larson, Christopher G. Chute, Gail P. Jarvik, James D. Ralston, Alan R. Shuldiner, Iftikhar J. Kullo, Peter Tarczy-Hornoch, Marc S. Williams |
AMIA | 11 |
| 2014 | Adverse Drug Event-based Stratification of Tumor Mutations: A Case Study of Breast Cancer Patients Receiving Aromatase Inhibitors
Michael T. Zimmermann, Naresh Prodduturi, Christopher G. Chute, Guoqian Jiang |
AMIA | 4 |
| 2014 | iGenetics: An Individualized Genetic Test Recommendation System Based on EHRs
Qian Zhu 0003, Christopher G. Chute, Matthew Ferber |
AMIA | 3 |
| 2014 | Genetic testing knowledge base (GTKB) towards individualized genetic test recommendation - An experimental studyabstractThe gap between a large growing number of genetic tests and a suboptimal clinical workflow of incorporating these tests into regular clinical practice poses barriers to effective reliance on advanced genetic technologies to improve quality of healthcare. A promising solution to encourage and assist physicians to incorporate genetic tests in their clinical practice is an intelligent genetic test recommendation system for 1) providing a comprehensive view of genetic tests as education resources; 2) recommending the most appropriate genetic tests to patients based on clinical evidence. In this paper, we introduce a genetic testing knowledge base, called GTKB, which was designed to support further individualized genetic test recommendation. More specifically, we extracted clinical characteristics identified from Electronic Health Records (EHRs) that have been used as phenotypic information for linked archived biological material to accelerate research in individualized medicine, and well-documented public genetic testing resources including Genetic Testing Registry (GTR) and published genetic testing guidelines (GTG) to construct a genetic test orientated knowledge base, ultimately supporting genetic test recommendation. An experimental study for “wilson disease mutation screen test” has been conducted to demonstrate the identification of salient clinical characteristics and the process of incorporating EHR derived phenotypes into the GTKB construction. Qian Zhu 0003, Christopher G. Chute, Matthew Ferber |
BIBM | 3 |
| 2014 | Health data use, stewardship, and governance: ongoing gaps and challenges: a report from AMIA's 2012 Health Policy MeetingabstractLarge amounts of personal health data are being collected and made available through existing and emerging technological media and tools. While use of these data has significant potential to facilitate research, improve quality of care for individuals and populations, and reduce healthcare costs, many policy-related issues must be addressed before their full value can be realized. These include the need for widely agreed-on data stewardship principles and effective approaches to reduce or eliminate data silos and protect patient privacy. AMIA's 2012 Health Policy Meeting brought together healthcare academics, policy makers, and system stakeholders (including representatives of patient groups) to consider these topics and formulate recommendations. A review of a set of Proposed Principles of Health Data Use led to a set of findings and recommendations, including the assertions that the use of health data should be viewed as a public good and that achieving the broad benefits of this use will require understanding and support from patients. George Hripcsak, Meryl Bloomrosen, Patricia Flatley Brennan, Christopher G. Chute, James J. Cimino, Don E. Detmer, Margo Edmunds, Peter J. Embí, Melissa M. Goldstein, William Edward Hammond, Gail M. Keenan, Steven E. Labkoff, Shawn P. Murphy, Charles Safran, Stuart M. Speedie, Howard R. Strasberg, Freda Temple, Adam B. Wilcox |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Research and applications: MedXN: an open source medication extraction and normalization tool for clinical textabstractOBJECTIVE: We developed the Medication Extraction and Normalization (MedXN) system to extract comprehensive medication information and normalize it to the most appropriate RxNorm concept unique identifier (RxCUI) as specifically as possible. METHODS: Medication descriptions in clinical notes were decomposed into medication name and attributes, which were separately extracted using RxNorm dictionary lookup and regular expression. Then, each medication name and its attributes were combined together according to RxNorm convention to find the most appropriate RxNorm representation. To do this, we employed serialized hierarchical steps implemented in Apache's Unstructured Information Management Architecture. We also performed synonym expansion, removed false medications, and employed inference rules to improve the medication extraction and normalization performance. RESULTS: An evaluation on test data of 397 medication mentions showed F-measures of 0.975 for medication name and over 0.90 for most attributes. The RxCUI assignment produced F-measures of 0.932 for medication name and 0.864 for full medication information. Most false negative RxCUI assignments in full medication information are due to human assumption of missing attributes and medication names in the gold standard. CONCLUSIONS: The MedXN system (http://sourceforge.net/projects/ohnlp/files/MedXN/) was able to extract comprehensive medication information with high accuracy and demonstrated good normalization capability to RxCUI as long as explicit evidence existed. More sophisticated inference rules might result in further improvements to specific RxCUI assignments for incomplete medication descriptions. Sunghwan Sohn, Cheryl Clark, Scott R. Halgrim, Sean P. Murphy, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | Building A Platform for Supporting Clinical Study Meta-Data Standards Authoring Using Scalable Semantic Web Technologies
Guoqian Jiang, Harold R. Solbrig, Christopher G. Chute |
AMIA | 3 |
| 2013 | Prioritize journals relevant to a clinical topic: A survey of US cardiologists about heart failure
Siddhartha Jonnalagadda, Christopher G. Chute |
AMIA | 2 |
| 2013 | Coreference Resolution from Medical Corpus with Topic Modeling
Dingcheng Li, Liwei Wang 0010, Cui Tao, Christopher G. Chute |
AMIA | 4 |
| 2013 | PhenotypePortal: An Open-Source Library and Platform for Authoring, Executing and Visualization of Electronic Health Records Driven Phenotyping Algorithms
Jyotishman Pathak, Cory M. Endle, Dale Suesse, Kevin J. Peterson, Craig Stancle, Dingcheng Li, Christopher G. Chute |
AMIA | 7 |
| 2013 | Mining Drug-Drug Interaction Patterns from Linked Clinical Data
Jyotishman Pathak, Richard C. Kiefer, Christopher G. Chute |
AMIA | 3 |
| 2013 | Medication Extraction and Normalization from Clinical Notes
Sunghwan Sohn, Cheryl Clark, Scott R. Halgrim, Sean P. Murphy, Christopher G. Chute |
AMIA | 5 |
| 2013 | Clinical Drug Extraction and Normalization from Clinical Notes
Sunghwan Sohn, Cheryl Clark, Scott R. Halgrim, Sean P. Murphy, Christopher G. Chute |
AMIA | 5 |
| 2013 | Psychological Assessment Instruments: A Coverage Analysis Using SNOMED CT, LOINC and QS Terminology
Piper A. Svensson-Ranallo, Terrence Adam, Katharine J. Nelson, Robert Krueger, Martin LaVenture, Christopher G. Chute |
AMIA | 6 |
| 2013 | Introduction and Implementation of Common Terminology Services 2 (CTS2)
Cui Tao, Harold R. Solbrig, Craig Stancle, Kevin J. Peterson, Cory M. Endle, Scott Bauer, Deepak K. Sharma, Christopher G. Chute |
AMIA | 8 |
| 2013 | Using standardized clinical data modeling and knowledge representation to compute pharmacogenomic data elements
Qian Zhu 0003, Jyotishman Pathak, Robert R. Freimuth, Christopher G. Chute |
AMIA | 4 |
| 2013 | Prioritizing journals relevant to a topic for addressing clinicians' information needsabstractObjective: Point of care access to knowledge from full text journal articles supports decision making and decreases medical errors. However, it is an overwhelming task to obtain full text for all journals and the quality of information for a specific topic such as Congestive Heart Failure (CHF) varies across the journals. In this paper, we develop a method to automatically rate journals for a given clinical topic to enable filtering journals or ranking the articles based on source journal. Materials and Methods: We surveyed 169 cardiologists chosen across the US that practice medicine and publish on the topic of CHF. They provided subjective opinion on how valuable the information provided in 60 journals (chosen by Mayo Clinic cardiologists) is for their clinical decision-making. Results: We identified the top CHF journals chosen by cardiologists and the objective metrics that correlate with their scores. Our best Multiple Linear Regression model has a correlation of 0.880 based on five-fold cross-validation. Discussion: We demonstrated that general journal metrics such as impact factor, h-index and number of articles per year provide better results when used in combination with topic-specific metrics such as number of abstracts indexed with the corresponding MeSH® terms. Conclusion: We obtained a journal priority score formula based on different journal metrics to automatically rate any journal in relation to its perceived importance in CHF. Our study shows that using this formula might yield a reliable list of prioritized journals even across other clinical topics such as Multiple Sclerosis. Siddhartha Jonnalagadda, Soheil Moosavinasab, Dingcheng Li, Martin D. Abel, Christopher G. Chute |
BIBM | 5 |
| 2013 | Towards assigning references using semantic, journal and citation relevanceabstractClinical knowledge systems (CKS) facilitate healthcare practitioners to make the best decisions at the point of care by having access to most recent knowledge. However, because of the overabundance of clinical research articles, it is time consuming to add new content manually to any CKS such as UpToDate©and to ensure that it is consistent with evidence. For this reason, many CKS such as Mayo Clinic's AskMayoExpert use expert-based (and not evidence-based) knowledge learned from peer-to-peer communication. The main motivation of this work is to explore methods for assigning references automatically to expert-written content. The system employs information retrieval (IR) techniques to retrieve related sentences from MEDLINE abstracts as evidence candidates and then uses a machine learning model trained with physician's evaluation score on certain journals for re-ranking. Finally, the system utilizes citation counts of individual articles for further refining the ranking. A preliminary evaluation using 59 UpToDate sentences and respective citations as gold standard showed that the median ranking improved 11 folds after adding the journal and citation relevance metrics on the top of baseline that only uses the semantic relevance metric. The system seems promising and ready for trials in real use-case scenarios with experts. Dingcheng Li, Christopher G. Chute, Siddhartha Jonnalagadda |
BIBM | 3 |
| 2013 | Mining drug-drug interaction patterns from linked data: A case study for Warfarin, Clopidogrel, and SimvastatinabstractBy nature, healthcare data is highly complex and voluminous. While on one hand, it provides unprecedented opportunities to identify hidden and unknown relationships between patients and treatment outcomes, or drugs and allergic reactions for given individuals, representing and querying large network datasets poses significant technical challenges. In this research, we study the use of Semantic Web and Linked Data technologies for identifying potential drug-drug interaction (DDI) information from publicly available resources, and determining if such interactions were observed using real patient data. Specifically, we apply Linked Data principles and technologies for representing patient data from electronic health records (EHRs) at Mayo Clinic as Resource Description Framework (RDF) graphs, and identify potential DDIs for three widely prescribed cardiovascular drugs: Warfarin, Clopidogrel and Simvastatin. Our results from the proof-of-concept study demonstrate the potential of applying such a methodology to study patient health outcomes as well as enabling genome-guided drug therapies and treatment interventions. Jyotishman Pathak, Richard C. Kiefer, Christopher G. Chute |
BIBM | 3 |
| 2013 | A semantic-web oriented representation of the clinical element model for secondary use of electronic health records dataabstractThe clinical element model (CEM) is an information model designed for representing clinical information in electronic health records (EHR) systems across organizations. The current representation of CEMs does not support formal semantic definitions and therefore it is not possible to perform reasoning and consistency checking on derived models. This paper introduces our efforts to represent the CEM specification using the Web Ontology Language (OWL). The CEM-OWL representation connects the CEM content with the Semantic Web environment, which provides authoring, reasoning, and querying tools. This work may also facilitate the harmonization of the CEMs with domain knowledge represented in terminology models as well as other clinical information models such as the openEHR archetype model. We have created the CEM-OWL meta ontology based on the CEM specification. A convertor has been implemented in Java to automatically translate detailed CEMs from XML to OWL. A panel evaluation has been conducted, and the results show that the OWL modeling can faithfully represent the CEM specification and represent patient data. Cui Tao, Guoqian Jiang, Thomas A. Oniki, Robert R. Freimuth, Qian Zhu 0003, Deepak K. Sharma, Jyotishman Pathak, Stanley M. Huff, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 9 |
| 2013 | Terminology representation guidelines for biomedical ontologies in the semantic web notations
Cui Tao, Jyotishman Pathak, Harold R. Solbrig, Wei-Qi Wei, Christopher G. Chute |
J. Biomed. Informatics | 5 |
| 2013 | Semantator: Semantic annotator for converting biomedical text to linked data
Cui Tao, Dezhao Song, Deepak K. Sharma, Christopher G. Chute |
J. Biomed. Informatics | 4 |
| 2013 | Harmonization and semantic annotation of data dictionaries from the Pharmacogenomics Research Network: A case study
Qian Zhu 0003, Robert R. Freimuth, Zonghui Lian, Scott Bauer, Jyotishman Pathak, Cui Tao, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 8 |
| 2013 | Disambiguation of PharmGKB drug-disease relations with NDF-RT and SPL
Qian Zhu 0003, Robert R. Freimuth, Jyotishman Pathak, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 5 |
| 2012 | Using PheWAS to Assess Pleiotropy of Genetic Risk Scores for Rheumatoid Arthritis and Coronary Artery Disease in the eMERGE Network
Robert J. Carroll, Katherine P. Liao, Anne E. Eyler, Lisa Bastarache, Dana C. Crawford, Peggy L. Peissig, Jyotishman Pathak, David Carrell, Abel N. Kho, Rongling Li, Daniel R. Masys, Gail P. Jarvik, Christopher G. Chute, Rex L. Chisholm, Eric B. Larson, Catherine A. McCarty, Iftikhar J. Kullo |
AMIA | 13 |
| 2012 | Comprehensive, Population-based, Peer-to-Peer Health Information Exchange and Clinical Data Repository Creation among Rural Practices and County Public Health Departments: The SE MN Beacon Community
Christopher G. Chute, Alex Alexander, Timothy Peters, Kris Riess, Rod Hughbanks, Lacey Hart, Larry Lemmon, Danial Jensen, Calvin E. Beebe |
AMIA | 1 |
| 2012 | Visualization and Reporting of Results for Electronic Health Records Driven Phenotyping using the Open-Source popHealth Platform
Cory M. Endle, Sahana Murthy, Dale Suesse, Craig Stancle, Dingcheng Li, Lacey Hart, Christopher G. Chute, Jyotishman Pathak |
AMIA | 7 |
| 2012 | The Pan-SHARP Project: An Interdisciplinary approach to Medication Reconciliation
Jorge R. Herskovic, Christopher G. Chute, Julian M. Goldman, Bernie Ács, David A. Kreda |
AMIA | 2 |
| 2012 | A Proposal Provenance Model for ICD-11 Revision Beta Phase
Guoqian Jiang, Cory M. Endle, Harold R. Solbrig, Christopher G. Chute |
AMIA | 4 |
| 2012 | Modeling and Executing Electronic Health Records Driven Phenotyping Algorithms using the NQF Quality Data Model and JBoss® Drools Engine
Dingcheng Li, Sahana Murthy, Davide Sottara, Christopher G. Chute, Stanley M. Huff, Jyotishman Pathak, Cory M. Endle, Dale Suesse, Craig Stancle |
AMIA | 4 |
| 2012 | Towards a semantic lexicon for clinical natural language processing
Stephen T. Wu, Dingcheng Li, Siddhartha Jonnalagadda, Sunghwan Sohn, Kavishwar B. Wagholikar, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
AMIA | 9 |
| 2012 | Using Electronic Health Records to Identify Patient Cohorts for Drug-Induced Thrombocytopenia, Neutropenia and Liver Injury
Jyotishman Pathak, Aref Al-Kali, Jayant Talwalkar, Abel N. Kho, Joshua C. Denny, Sean P. Murphy, Kevin Bruce, Matthew J. Durski, Christopher G. Chute |
AMIA | 9 |
| 2012 | Mining the Human Phenome using Semantic Web Technologies: A Case Study for Type 2 Diabetes
Jyotishman Pathak, Richard C. Kiefer, Suzette J. Bielinski, Christopher G. Chute |
AMIA | 4 |
| 2012 | Mining Genotype-Phenotype Associations from Electronic Health Records and Biorepositories using Semantic Web Technologies
Jyotishman Pathak, Richard C. Kiefer, Robert R. Freimuth, Suzette J. Bielinski, Christopher G. Chute |
AMIA | 5 |
| 2012 | Common Terminology Services 2 (CTS2) for Biomedical Community
Cui Tao, Harold R. Solbrig, Pradip Kanjamala, Kevin J. Peterson, Craig Stancle, Christopher G. Chute |
AMIA | 6 |
| 2012 | Tracking Immigrant Health with Natural Language Processing
Stephen T. Wu, Mark Wieland, Vinod Kaggal, Christopher G. Chute |
AMIA | 5 |
| 2012 | A Standardized Drug and Drug Class Universal Network
Qian Zhu 0003, Guoqian Jiang, Christopher G. Chute |
AMIA | 3 |
| 2012 | (1) Obstacles and options for big-data applications in biomedicine: The role of standards and normalizationsabstractAdvances in computing capabilities are palpably evident throughout many industries manifest by unprecedented, large-scale data integration and inferencing. Branded as "big-data" in many cases, the question of whether such techniques can leverage advances in biomedicine and clinical practice are obvious. High-throughput clinical analytics, synthesizing genomic and clinical attributes of a particular patient, portends predictive models that can directly influence clinical care decisions. However, to make this widely shared vision practical and scalable, barriers attributable to data heterogeneity dominate. Methods and strategies to increase the comparability and consistency of healthcare related data will be discussed. Christopher G. Chute |
BIBM | 1 |
| 2012 | Quality evaluation of value sets from cancer study common data elements using the UMLS semantic groupsabstractOBJECTIVE: The objective of this study is to develop an approach to evaluate the quality of terminological annotations on the value set (ie, enumerated value domain) components of the common data elements (CDEs) in the context of clinical research using both unified medical language system (UMLS) semantic types and groups. MATERIALS AND METHODS: The CDEs of the National Cancer Institute (NCI) Cancer Data Standards Repository, the NCI Thesaurus (NCIt) concepts and the UMLS semantic network were integrated using a semantic web-based framework for a SPARQL-enabled evaluation. First, the set of CDE-permissible values with corresponding meanings in external controlled terminologies were isolated. The corresponding value meanings were then evaluated against their NCI- or UMLS-generated semantic network mapping to determine whether all of the meanings fell within the same semantic group. RESULTS: Of the enumerated CDEs in the Cancer Data Standards Repository, 3093 (26.2%) had elements drawn from more than one UMLS semantic group. A random sample (n=100) of this set of elements indicated that 17% of them were likely to have been misclassified. DISCUSSION: The use of existing semantic web tools can support a high-throughput mechanism for evaluating the quality of large CDE collections. This study demonstrates that the involvement of multiple semantic groups in an enumerated value domain of a CDE is an effective anchor to trigger an auditing point for quality evaluation activities. CONCLUSION: This approach produces a useful quality assurance mechanism for a clinical study CDE repository. Guoqian Jiang, Harold R. Solbrig, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Use of diverse electronic medical record systems to identify genetic risk for type 2 diabetes within a genome-wide association studyabstractOBJECTIVE: Genome-wide association studies (GWAS) require high specificity and large numbers of subjects to identify genotype-phenotype correlations accurately. The aim of this study was to identify type 2 diabetes (T2D) cases and controls for a GWAS, using data captured through routine clinical care across five institutions using different electronic medical record (EMR) systems. MATERIALS AND METHODS: An algorithm was developed to identify T2D cases and controls based on a combination of diagnoses, medications, and laboratory results. The performance of the algorithm was validated at three of the five participating institutions compared against clinician review. A GWAS was subsequently performed using cases and controls identified by the algorithm, with samples pooled across all five institutions. RESULTS: The algorithm achieved 98% and 100% positive predictive values for the identification of diabetic cases and controls, respectively, as compared against clinician review. By standardizing and applying the algorithm across institutions, 3353 cases and 3352 controls were identified. Subsequent GWAS using data from five institutions replicated the TCF7L2 gene variant (rs7903146) previously associated with T2D. DISCUSSION: By applying stringent criteria to EMR data collected through routine clinical care, cases and controls for a GWAS were identified that subsequently replicated a known genetic variant. The use of standard terminologies to define data elements enabled pooling of subjects and data across five different institutions to achieve the robust numbers required for GWAS. CONCLUSIONS: An algorithm using commonly available data from five different EMR can accurately identify T2D cases and controls for genetic study across multiple institutions. Abel N. Kho, M. Geoffrey Hayes, Laura Rasmussen-Torvik, Jennifer A. Pacheco, William K. Thompson, Loren L. Armstrong, Joshua C. Denny, Peggy L. Peissig, Aaron W. Miller, Wei-Qi Wei, Suzette J. Bielinski, Christopher G. Chute, Cynthia L. Leibson, Gail P. Jarvik, David R. Crosslin, Christopher S. Carlson, Katherine M. Newton, Wendy A. Wolf, Rex L. Chisholm, William L. Lowe |
J. Am. Medical Informatics Assoc. | 12 |
| 2012 | The National Center for Biomedical OntologyabstractThe National Center for Biomedical Ontology is now in its seventh year. The goals of this National Center for Biomedical Computing are to: create and maintain a repository of biomedical ontologies and terminologies; build tools and web services to enable the use of ontologies and terminologies in clinical and translational research; educate their trainees and the scientific community broadly about biomedical ontology and ontology-based technology and best practices; and collaborate with a variety of groups who develop and use ontologies and terminologies in biomedicine. The centerpiece of the National Center for Biomedical Ontology is a web-based resource known as BioPortal. BioPortal makes available for research in computationally useful forms more than 270 of the world's biomedical ontologies and terminologies, and supports a wide range of web services that enable investigators to use the ontologies to annotate and retrieve data, to generate value sets and special-purpose lexicons, and to perform advanced analytics on a wide range of biomedical data. Mark A. Musen, Natasha F. Noy, Nigam H. Shah, Patricia L. Whetzel, Christopher G. Chute, Margaret-Anne D. Storey, Barry Smith 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Automated discovery of drug treatment patterns for endocrine therapy of breast cancer within an electronic medical recordabstractOBJECTIVE: To develop an algorithm for the discovery of drug treatment patterns for endocrine breast cancer therapy within an electronic medical record and to test the hypothesis that information extracted using it is comparable to the information found by traditional methods. MATERIALS: The electronic medical charts of 1507 patients diagnosed with histologically confirmed primary invasive breast cancer. METHODS: The automatic drug treatment classification tool consisted of components for: (1) extraction of drug treatment-relevant information from clinical narratives using natural language processing (clinical Text Analysis and Knowledge Extraction System); (2) extraction of drug treatment data from an electronic prescribing system; (3) merging information to create a patient treatment timeline; and (4) final classification logic. RESULTS: Agreement between results from the algorithm and from a nurse abstractor is measured for categories: (0) no tamoxifen or aromatase inhibitor (AI) treatment; (1) tamoxifen only; (2) AI only; (3) tamoxifen before AI; (4) AI before tamoxifen; (5) multiple AIs and tamoxifen cycles in no specific order; and (6) no specific treatment dates. Specificity (all categories): 96.14%-100%; sensitivity (categories (0)-(4)): 90.27%-99.83%; sensitivity (categories (5)-(6)): 0-23.53%; positive predictive values: 80%-97.38%; negative predictive values: 96.91%-99.93%. DISCUSSION: Our approach illustrates a secondary use of the electronic medical record. The main challenge is event temporality. CONCLUSION: We present an algorithm for automated treatment classification within an electronic medical record to combine information extracted through natural language processing with that extracted from structured databases. The algorithm has high specificity for all categories, high sensitivity for five categories, and low sensitivity for two categories. Guergana K. Savova, Janet E. Olson, Sean P. Murphy, Victoria L. Cafourek, Fergus J. Couch, Matthew P. Goetz, James N. Ingle, Vera J. Suman, Christopher G. Chute, Richard M. Weinshilboum |
J. Am. Medical Informatics Assoc. | 9 |
| 2012 | Impact of data fragmentation across healthcare centers on the accuracy of a high-throughput clinical phenotyping algorithm for specifying subjects with type 2 diabetes mellitusabstractOBJECTIVE: To evaluate data fragmentation across healthcare centers with regard to the accuracy of a high-throughput clinical phenotyping (HTCP) algorithm developed to differentiate (1) patients with type 2 diabetes mellitus (T2DM) and (2) patients with no diabetes. MATERIALS AND METHODS: This population-based study identified all Olmsted County, Minnesota residents in 2007. We used provider-linked electronic medical record data from the two healthcare centers that provide >95% of all care to County residents (ie, Olmsted Medical Center and Mayo Clinic in Rochester, Minnesota, USA). Subjects were limited to residents with one or more encounter January 1, 2006 through December 31, 2007 at both healthcare centers. DM-relevant data on diagnoses, laboratory results, and medication from both centers were obtained during this period. The algorithm was first executed using data from both centers (ie, the gold standard) and then from Mayo Clinic alone. Positive predictive values and false-negative rates were calculated, and the McNemar test was used to compare categorization when data from the Mayo Clinic alone were used with the gold standard. Age and sex were compared between true-positive and false-negative subjects with T2DM. Statistical significance was accepted as p<0.05. RESULTS: With data from both medical centers, 765 subjects with T2DM (4256 non-DM subjects) were identified. When single-center data were used, 252 T2DM subjects (1573 non-DM subjects) were missed; an additional false-positive 27 T2DM subjects (215 non-DM subjects) were identified. The positive predictive values and false-negative rates were 95.0% (513/540) and 32.9% (252/765), respectively, for T2DM subjects and 92.6% (2683/2898) and 37.0% (1573/4256), respectively, for non-DM subjects. Age and sex distribution differed between true-positive (mean age 62.1; 45% female) and false-negative (mean age 65.0; 56.0% female) T2DM subjects. CONCLUSION: The findings show that application of an HTCP algorithm using data from a single medical center contributes to misclassification. These findings should be considered carefully by researchers when developing and executing HTCP algorithms. Wei-Qi Wei, Cynthia L. Leibson, Jeanine E. Ransom, Abel N. Kho, Pedro J. Caraballo, High Seng Chai, Barbara P. Yawn, Jennifer A. Pacheco, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 9 |
| 2012 | Unified Medical Language System term occurrences in clinical notes: a large-scale corpus analysisabstractOBJECTIVE: To characterise empirical instances of Unified Medical Language System (UMLS) Metathesaurus term strings in a large clinical corpus, and to illustrate what types of term characteristics are generalisable across data sources. DESIGN: Based on the occurrences of UMLS terms in a 51 million document corpus of Mayo Clinic clinical notes, this study computes statistics about the terms' string attributes, source terminologies, semantic types and syntactic categories. Term occurrences in 2010 i2b2/VA text were also mapped; eight example filters were designed from the Mayo-based statistics and applied to i2b2/VA data. RESULTS: For the corpus analysis, negligible numbers of mapped terms in the Mayo corpus had over six words or 55 characters. Of source terminologies in the UMLS, the Consumer Health Vocabulary and Systematized Nomenclature of Medicine-Clinical Terms (SNOMED-CT) had the best coverage in Mayo clinical notes at 106426 and 94788 unique terms, respectively. Of 15 semantic groups in the UMLS, seven groups accounted for 92.08% of term occurrences in Mayo data. Syntactically, over 90% of matched terms were in noun phrases. For the cross-institutional analysis, using five example filters on i2b2/VA data reduces the actual lexicon to 19.13% of the size of the UMLS and only sees a 2% reduction in matched terms. CONCLUSION: The corpus statistics presented here are instructive for building lexicons from the UMLS. Features intrinsic to Metathesaurus terms (well formedness, length and language) generalise easily across clinical institutions, but term frequencies should be adapted with caution. The semantic groups of mapped terms may differ slightly from institution to institution, but they differ greatly when moving to the biomedical literature domain. Stephen T. Wu, Dingcheng Li, Cui Tao, Mark A. Musen, Christopher G. Chute, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 6 |
| 2012 | Building a robust, scalable and standards-driven infrastructure for secondary use of EHR data: The SHARPn project
Susan Rea, Jyotishman Pathak, Guergana K. Savova, Thomas A. Oniki, Les Westberg, Calvin E. Beebe, Cui Tao, Craig G. Parker, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
J. Biomed. Informatics | 11 |
| 2012 | Cross-terminology mapping challenges: A demonstration using medication terminological systems
Himali Saitwal, David Qing, Elmer V. Bernstam, Christopher G. Chute, Todd R. Johnson |
J. Biomed. Informatics | 5 |
| 2012 | Erratum to "Cross-terminology mapping challenges: A demonstration using medication terminological systems" [J. Biomed. Inform. (2012) 613-625]
Himali Saitwal, David Qing, Elmer V. Bernstam, Christopher G. Chute, Todd R. Johnson |
J. Biomed. Informatics | 5 |
| 2011 | Letter: Further revamping VA's NDF-RT drug terminology for clinical researchabstractBiomedical terminology and vocabulary standards (for the purposes of this correspondence, we use the terms ‘terminology,’ ‘vocabulary’ and ‘ontology’ interchangeably.) play an important role in enabling consistent, comparable, and meaningful sharing of data within and across institutional boundaries, as well as ensuring semantic interoperability. An important domain for developing standardized vocabularies is medications, where existing standards structure and organize approved drug products and ingredients by various characteristics or properties to support a multitude of clinical and epidemiological research questions across the spectrum of health and disease. Veteran Affairs' National Drug File-Reference Terminology (NDF-RT; see figure 1) is a Federal Medication-recommended standardized terminology resource encompassing medications, ingredients, and high-level drug classes for Chemical Structure (eg, Acetanilides), Mechanism of Action (eg, Prostaglandin Receptor Antagonists), Physiological Effect (eg, Decreased Prostaglandin Production), drug–disease relationship describing the Therapeutic Intent (eg, Pain), and Pharmacokinetics describing the mechanisms of absorption and distribution of an administered drug within a body (eg, Hepatic Metabolism). Additionally, NDF-RT contains two independent lists of drug classes: Legacy VA classes and External Pharmacologic classes, where the former simply provides a shallow hierarchy of ‘clinically oriented’ classes (eg, β blockers), and the latter focuses on classifying drugs based on their chemical properties and functional groups (eg, H1 Receptor Antagonist). National Drug File-Reference Terminology drug-class hierarchy organization (adapted from Carter et al4). In the recent past, several research reports have highlighted important issues and challenges in using NDF-RT for clinical research and interoperability. Bodenreider et al1 focused on determining anticoagulation status of patients based on a list of medications prescribed using NDF-RT as the underlying drug-class terminology. In particular, this work concentrated on leveraging description logics (DL)-based representation of NDF-RT to infer additional information about drug-class relationships using the Legacy VA classes and External Pharmacologic classes. During this process, the authors not only had to make significant modifications and re-engineering to NDF-RT's DL representation, but also encountered several missing pieces of drug-class membership information in NDF-RT. In another study, Palchuk et al2 constructed a hierarchy of NDF-RT drug classes with drug and medication information from another standardized drug terminology, RxNorm, using data from the patient's electronic medical record. Similar to Bodenreider et al,1 here authors had to perform significant re-engineering to map, and subsequently classify, RxNorm drug products using Legacy VA classes from NDF-RT. The authors found this process to be extremely onerous, and proposed the evolution of RxNorm toward an interface terminology with hierarchical and categorical organization. In our own work published in JAMIA,3 we investigated similar issues in mapping and classifying drugs and medication products from RxNorm using NDF-RT's multiaxial classification. We found several issues where the mappings were incomplete and, in many occasions, semantically and clinically inconsistent. Based on these recent findings, it has become abundantly evident that to leverage NDF-RT continually for clinical and epidemiological research, it is vital to address the existing issues. In particular, we highlight the following issues for consideration: Alignment with RxNorm: Palchuk et al2 and our previous work3 illustrated several problematic examples where the relationships between NDF-RT and RxNorm drug concept entities were either missing or misrepresented due to curation problems. Given that RxNorm does not currently provide hierarchical classification of drug products, the mappings between drug concepts in RxNorm and the corresponding classes in NDF-RT are vital for research projects that use RxNorm for coding their medication data. Missing and inconsistent mappings can lead to incorrect conclusions. Relationships between drug products and Legacy VA classes: The Legacy VA classes that were derived from the VA-NDF,4 while deprecated, still continue to provide significant value and merit with respect to ‘clinically relevant’ drug classification. However, in its current formalism, NDF-RT only allows assignment of a single Legacy VA class to a particular drug product. For example, even though it is clinically appropriate to classify a drug as both an antihypertensive and a β-blocker, in reality a majority of drug products in NDF-RT are assigned a single Legacy VA class. We believe that this limitation needs to be addressed, since the Legacy VA classes, although derived from the legacy VA-NDF, have significant clinical implications for drug classifications. Relationships between drug products and External Pharmacologic classes: As illustrated by Bodenreider et al,1 NDF-RT in its current form does not contain any relationships between ingredients and External Pharmacologic classes. As an example, in NDF-RT, Clopidogrel and Platelet Aggregation Inhibitor are not related via any relationship—either direct or indirect. Arguably, this is a significant limitation and has implications with respect to drug classifications and querying. Metadata annotations: Finally, we believe that NDF-RT should follow best practices for vocabulary and terminology development.5 In particular, what is notably missing from NDF-RT are appropriate metadata annotations for different drug, ingredient, and drug-class entities. As an example, the Legacy VA class ‘Loop Diuretics’ and External Pharmacologic class ‘Loop Diuretic’ are distinguished only by a slight difference in the label name, without additional annotation indicating their differences, similarities, etc. Consequently, someone unfamiliar with NDF-RT multiaxial classification runs the risk of using the incorrect classification for her application. In summary, we hope that, via this correspondence, we have highlighted some of the important issues with NDF-RT, which arguably is emerging as one of the most important standardized public drug-classification terminologies. Our expectation is that addressing the above problems in future NDF-RT releases will significantly benefit the clinical research informatics community. This work was supported by National Human Genome Research Institute. None. Not commissioned; externally peer reviewed. Jyotishman Pathak, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Mapping clinical phenotype data elements to standardized metadata repositories and controlled terminologies: the eMERGE Network experienceabstractBACKGROUND: Systematic study of clinical phenotypes is important for a better understanding of the genetic basis of human diseases and more effective gene-based disease management. A key aspect in facilitating such studies requires standardized representation of the phenotype data using common data elements (CDEs) and controlled biomedical vocabularies. In this study, the authors analyzed how a limited subset of phenotypic data is amenable to common definition and standardized collection, as well as how their adoption in large-scale epidemiological and genome-wide studies can significantly facilitate cross-study analysis. METHODS: The authors mapped phenotype data dictionaries from five different eMERGE (Electronic Medical Records and Genomics) Network sites studying multiple diseases such as peripheral arterial disease and type 2 diabetes. For mapping, standardized terminological and metadata repository resources, such as the caDSR (Cancer Data Standards Registry and Repository) and SNOMED CT (Systematized Nomenclature of Medicine), were used. The mapping process comprised both lexical (via searching for relevant pre-coordinated concepts and data elements) and semantic (via post-coordination) techniques. Where feasible, new data elements were curated to enhance the coverage during mapping. A web-based application was also developed to uniformly represent and query the mapped data elements from different eMERGE studies. RESULTS: Approximately 60% of the target data elements (95 out of 157) could be mapped using simple lexical analysis techniques on pre-coordinated terms and concepts before any additional curation of terminology and metadata resources was initiated by eMERGE investigators. After curation of 54 new caDSR CDEs and nine new NCI thesaurus concepts and using post-coordination, the authors were able to map the remaining 40% of data elements to caDSR and SNOMED CT. A web-based tool was also implemented to assist in semi-automatic mapping of data elements. CONCLUSION: This study emphasizes the requirement for standardized representation of clinical research data using existing metadata and terminology resources and provides simple techniques and software for data element mapping using experiences from the eMERGE Network. Jyotishman Pathak, Janey Wang, Sudha Kashyap, Melissa A. Basford, Rongling Li, Daniel R. Masys, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 7 |
| 2011 | Drug side effect extraction from clinical narratives of psychiatry and psychology patientsabstractOBJECTIVE: To extract physician-asserted drug side effects from electronic medical record clinical narratives. MATERIALS AND METHODS: Pattern matching rules were manually developed through examining keywords and expression patterns of side effects to discover an individual side effect and causative drug relationship. A combination of machine learning (C4.5) using side effect keyword features and pattern matching rules was used to extract sentences that contain side effect and causative drug pairs, enabling the system to discover most side effect occurrences. Our system was implemented as a module within the clinical Text Analysis and Knowledge Extraction System. RESULTS: The system was tested in the domain of psychiatry and psychology. The rule-based system extracting side effects and causative drugs produced an F score of 0.80 (0.55 excluding allergy section). The hybrid system identifying side effect sentences had an F score of 0.75 (0.56 excluding allergy section) but covered more side effect and causative drug pairs than individual side effect extraction. DISCUSSION: The rule-based system was able to identify most side effects expressed by clear indication words. More sophisticated semantic processing is required to handle complex side effect descriptions in the narrative. We demonstrated that our system can be trained to identify sentences with complex side effect descriptions that can be submitted to a human expert for further abstraction. CONCLUSION: Our system was able to extract most physician-asserted drug side effects. It can be used in either an automated mode for side effect extraction or semi-automated mode to identify side effect sentences that can significantly simplify abstraction by a human expert. Sunghwan Sohn, Jean-Pierre A. Kocher, Christopher G. Chute, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | Quality evaluation of cancer study Common Data Elements using the UMLS Semantic NetworkabstractThe binding of controlled terminology has been regarded as important for standardization of Common Data Elements (CDEs) in cancer research. However, the potential of such binding has not yet been fully explored, especially its quality assurance aspect. The objective of this study is to explore whether there is a relationship between terminological annotations and the UMLS Semantic Network (SN) that can be exploited to improve those annotations. We profiled the terminological concepts associated with the standard structure of the CDEs of the NCI Cancer Data Standards Repository (caDSR) using the UMLS SN. We processed 17798 data elements and extracted 17526 primary object class/property concept pairs. We identified dominant semantic types for the categories "object class" and "property" and determined that the preponderance of the instances were disjoint (i.e. the intersection of semantic types between the two categories is empty). We then performed a preliminary evaluation on the data elements whose asserted primary object class/property concept pairs conflict with this observation - where the semantic type of the object class fell into a SN category typically used by property or visa-versa. In conclusion, the UMLS SN based profiling approach is feasible for the quality assurance and accessibility of the cancer study CDEs. This approach could provide useful insight about how to build mechanisms of quality assurance in a meta-data repository. Guoqian Jiang, Harold R. Solbrig, Christopher G. Chute |
J. Biomed. Informatics | 3 |
| 2011 | Towards a framework for developing semantic relatedness reference standards
Serguei V. S. Pakhomov, Ted Pedersen, Bridget T. McInnes, Genevieve B. Melton, Alexander Ruggieri, Christopher G. Chute |
J. Biomed. Informatics | 6 |
| 2010 | Time-Oriented Question Answering from Clinical Narratives Using Semantic-Web Techniques
Cui Tao, Harold R. Solbrig, Deepak K. Sharma, Wei-Qi Wei, Guergana K. Savova, Christopher G. Chute |
ISWC (2) | 6 |
| 2010 | The Enterprise Data Trust at Mayo Clinic: a semantically integrated warehouse of biomedical dataabstractMayo Clinic's Enterprise Data Trust is a collection of data from patient care, education, research, and administrative transactional systems, organized to support information retrieval, business intelligence, and high-level decision making. Structurally it is a top-down, subject-oriented, integrated, time-variant, and non-volatile collection of data in support of Mayo Clinic's analytic and decision-making processes. It is an interconnected piece of Mayo Clinic's Enterprise Information Management initiative, which also includes Data Governance, Enterprise Data Modeling, the Enterprise Vocabulary System, and Metadata Management. These resources enable unprecedented organization of enterprise information about patient, genomic, and research data. While facile access for cohort definition or aggregate retrieval is supported, a high level of security, retrieval audit, and user authentication ensures privacy, confidentiality, and respect for the trust imparted by our patients for the respectful use of information about their conditions. Christopher G. Chute, Scott A. Beck, Thomas B. Fisk, David N. Mohr |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Leveraging informatics for genetic studies: use of the electronic medical record to enable a genome-wide association study of peripheral arterial diseaseabstractBACKGROUND: There is significant interest in leveraging the electronic medical record (EMR) to conduct genome-wide association studies (GWAS). METHODS: A biorepository of DNA and plasma was created by recruiting patients referred for non-invasive lower extremity arterial evaluation or stress ECG. Peripheral arterial disease (PAD) was defined as a resting/post-exercise ankle-brachial index (ABI) less than or equal to 0.9, a history of lower extremity revascularization, or having poorly compressible leg arteries. Controls were patients without evidence of PAD. Demographic data and laboratory values were extracted from the EMR. Medication use and smoking status were established by natural language processing of clinical notes. Other risk factors and comorbidities were ascertained based on ICD-9-CM codes, medication use and laboratory data. RESULTS: Of 1802 patients with an abnormal ABI, 115 had non-atherosclerotic vascular disease such as vasculitis, Buerger's disease, trauma and embolism (phenocopies) based on ICD-9-CM diagnosis codes and were excluded. The PAD cases (66+/-11 years, 64% men) were older than controls (61+/-8 years, 60% men) but had similar geographical distribution and ethnic composition. Among PAD cases, 1444 (85.6%) had an abnormal ABI, 233 (13.8%) had poorly compressible arteries and 10 (0.6%) had a history of lower extremity revascularization. In a random sample of 95 cases and 100 controls, risk factors and comorbidities ascertained from EMR-based algorithms had good concordance compared with manual record review; the precision ranged from 67% to 100% and recall from 84% to 100%. CONCLUSION: This study demonstrates use of the EMR to ascertain phenocopies, phenotype heterogeneity and relevant covariates to enable a GWAS of PAD. Biorepositories linked to EMR may provide a relatively efficient means of conducting GWAS. Iftikhar J. Kullo, Jyotishman Pathak, Guergana K. Savova, Zeenat Ali, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 6 |
| 2010 | Analyzing categorical information in two publicly available drug terminologies: RxNorm and NDF-RTabstractBACKGROUND: The RxNorm and NDF-RT (National Drug File Reference Terminology) are a suite of terminology standards for clinical drugs designated for use in the US federal government systems for electronic exchange of clinical health information. Analyzing how different drug products described in these terminologies are categorized into drug classes will help in their better organization and classification of pharmaceutical information. METHODS: Mappings between drug products in RxNorm and NDF-RT drug classes were extracted. Mappings were also extracted between drug products in RxNorm to five high-level NDF-RT categories: Chemical Structure; cellular or subcellular Mechanism of Action; organ-level or system-level Physiologic Effect; Therapeutic Intent; and Pharmacokinetics. Coverage for the mappings and the gaps were evaluated and analyzed algorithmically. RESULTS: Approximately 54% of RxNorm drug products (Semantic Clinical Drugs) were found not to have a correspondence in NDF-RT. Similarly, approximately 45% of drug products in NDF-RT are missing from RxNorm, most of which can be attributed to differences in dosage, strength, and route form. Approximately 81% of Chemical Structure classes, 42% of Mechanism of Action classes, 75% of Physiologic Effect classes, 76% of Therapeutic Intent classes, and 88% of Pharmacokinetics classes were also found not to have any RxNorm drug products classified under them. Finally, various issues regarding inconsistent mappings between drug concepts were identified in both terminologies. CONCLUSION: This investigation identified potential limitations of the existing classification systems and various issues in specification of correspondences between the concepts in RxNorm and NDF-RT. These proposals and methods provide the preliminary steps in addressing some of the requirements. Jyotishman Pathak, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Comparing and evaluating terminology services application programming interfaces: RxNav, UMLSKS and LexBIGabstractTo facilitate the integration of terminologies into applications, various terminology services application programming interfaces (API) have been developed in the recent past. In this study, three publicly available terminology services API, RxNav, UMLSKS and LexBIG, are compared and functionally evaluated with respect to the retrieval of information from one biomedical terminology, RxNorm, to which all three services provide access. A list of queries is established covering a wide spectrum of terminology services functionalities such as finding RxNorm concepts by their name, or navigating different types of relationships. Test data were generated from the RxNorm dataset to evaluate the implementation of the functionalities in the three API. The results revealed issues with various aspects of the API implementation (eg, handling of obsolete terms by LexBIG) and documentation (eg, navigational paths used in RxNav) that were subsequently addressed by the development teams of the three API investigated. Knowledge about such discrepancies helps inform the choice of an API for a given use case. Jyotishman Pathak, Lee B. Peters, Christopher G. Chute, Olivier Bodenreider |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applicationsabstractWe aim to build and evaluate an open-source natural language processing system for information extraction from electronic medical record clinical free-text. We describe and evaluate our system, the clinical Text Analysis and Knowledge Extraction System (cTAKES), released open-source at http://www.ohnlp.org. The cTAKES builds on existing open-source technologies-the Unstructured Information Management Architecture framework and OpenNLP natural language processing toolkit. Its components, specifically trained for the clinical domain, create rich linguistic and semantic annotations. Performance of individual components: sentence boundary detector accuracy=0.949; tokenizer accuracy=0.949; part-of-speech tagger accuracy=0.936; shallow parser F-score=0.924; named entity recognizer and system-level evaluation F-score=0.715 for exact and 0.824 for overlapping spans, and accuracy for concept mapping, negation, and status attributes for exact and overlapping spans of 0.957, 0.943, 0.859, and 0.580, 0.939, and 0.839, respectively. Overall performance is discussed against five applications. The cTAKES annotations are the foundation for methods and modules for higher-level semantic processing of clinical free-text. Guergana K. Savova, James J. Masanz, Philip V. Ogren, Jiaping Zheng, Sunghwan Sohn, Karin Kipper Schuler, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 7 |
| 2010 | An analytical approach to characterize morbidity profile dissimilarity between distinct cohorts using electronic medical records
Jonathan S. Schildcrout, Melissa A. Basford, Jill M. Pulley, Daniel R. Masys, Dan M. Roden, Deede Wang, Christopher G. Chute, Iftikhar J. Kullo, David Carrell, Peggy L. Peissig, Abel N. Kho, Joshua C. Denny |
J. Biomed. Informatics | 7 |
| 2009 | Viewpoint Paper: Auditing the Semantic Completeness of SNOMED CT Using Formal Concept AnalysisabstractOBJECTIVE: This study sought to develop and evaluate an approach for auditing the semantic completeness of the SNOMED CT contents using a formal concept analysis (FCA)-based model. DESIGN: We developed a model for formalizing the normal forms of SNOMED CT expressions using FCA. Anonymous nodes, identified through the analyses, were retrieved from the model for evaluation. Two quasi-Poisson regression models were developed to test whether anonymous nodes can evaluate the semantic completeness of SNOMED CT contents (Model 1), and for testing whether such completeness differs between 2 clinical domains (Model 2). The data were randomly sampled from all the contexts that could be formed in the 2 largest domains: Procedure and Clinical Finding. Case studies (n = 4) were performed on randomly selected anonymous node samples for validation. MEASUREMENTS: In Model 1, the outcome variable is the number of fully defined concepts within a context, while the explanatory variables are the number of lattice nodes and the number of anonymous nodes. In Model 2, the outcome variable is the number of anonymous nodes and the explanatory variables are the number of lattice nodes and a binary category for domain (Procedure/Clinical Finding). RESULTS: A total of 5,450 contexts from the 2 domains were collected for analyses. Our findings revealed that the number of anonymous nodes had a significant negative correlation with the number of fully defined concepts within a context (p < 0.001). Further, the Clinical Finding domain had fewer anonymous nodes than the Procedure domain (p < 0.001). Case studies demonstrated that the anonymous nodes are an effective index for auditing SNOMED CT. CONCLUSION: The anonymous nodes retrieved from FCA-based analyses are a candidate proxy for the semantic completeness of the SNOMED CT contents. Our novel FCA-based approach can be useful for auditing the semantic completeness of SNOMED CT contents, or any large ontology, within or across domains. Guoqian Jiang, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 2 |
| 2009 | Implementation Brief: LexGrid: A Framework for Representing, Storing, and Querying Biomedical Terminologies from Simple to SublimeabstractMany biomedical terminologies, classifications, and ontological resources such as the NCI Thesaurus (NCIT), International Classification of Diseases (ICD), Systematized Nomenclature of Medicine (SNOMED), Current Procedural Terminology (CPT), and Gene Ontology (GO) have been developed and used to build a variety of IT applications in biology, biomedicine, and health care settings. However, virtually all these resources involve incompatible formats, are based on different modeling languages, and lack appropriate tooling and programming interfaces (APIs) that hinder their wide-scale adoption and usage in a variety of application contexts. The Lexical Grid (LexGrid) project introduced in this paper is an ongoing community-driven initiative, coordinated by the Mayo Clinic Division of Biomedical Statistics and Informatics, designed to bridge this gap using a common terminology model called the LexGrid model. The key aspect of the model is to accommodate multiple vocabulary and ontology distribution formats and support of multiple data stores for federated vocabulary distribution. The model provides a foundation for building consistent and standardized APIs to access multiple vocabularies that support lexical search queries, hierarchy navigation, and a rich set of features such as recursive subsumption (e.g., get all the children of the concept penicillin). Existing LexGrid implementations include the LexBIG API as well as a reference implementation of the HL7 Common Terminology Services (CTS) specification providing programmatic access via Java, Web, and Grid services. Jyotishman Pathak, Harold R. Solbrig, James D. Buntrock, Thomas M. Johnson, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 5 |
| 2009 | Formalizing ICD coding rules using Formal Concept Analysis
Guoqian Jiang, Jyotishman Pathak, Christopher G. Chute |
J. Biomed. Informatics | 3 |
| 2008 | caBIG™ Compatibility Review System: Software to Support the Evaluation of Applications Using Defined Interoperability Criteria
Robert R. Freimuth, Michael W. Schauer, Preeti Lodha, Poornima Govindrao, Rakesh Nagarajan, Christopher G. Chute |
AMIA | 6 |
| 2008 | Designing and testing computer based screening engine for severe sepsis/septic shock
Vitaly Herasevich, Bekele Afessa, Christopher G. Chute, Ognjen Gajic |
AMIA | 3 |
| 2008 | LexValueSets: An Approach for Context-Driven Value Sets Extraction
Jyotishman Pathak, Guoqian Jiang, Sridhar O. Dwarkanath, James D. Buntrock, Christopher G. Chute |
AMIA | 5 |
| 2008 | Constructing Evaluation Corpora for Automated Clinical Named Entity Recognition
Philip V. Ogren, Guergana K. Savova, Christopher G. Chute |
LREC | 3 |
| 2008 | Editorial Comments: Biosurveillance, Classification, and Semantic Health TechnologiesabstractThe emergence of the World Wide Web in the early 1990s catapulted the Internet from a curiosity for computer scientists, engineers, and experimenters testing data interchange technologies into a vast, dynamic, and unprecedented information resource that has fundamentally transformed science, society, and human communications. The implications of this intellectual paradigm shift are still being explored, and the full impact of this dramatic transformation, which instantly makes available massive amounts of information, has not yet become fully appreciated. Nevertheless, the emerging dependency of the health sciences on increasingly practical semantic technologies to organize and leverage these vast information resources is now unquestioned. Ranging from the interpretive functionality of the human genome to public health surveillance among nations and continents, the full spectrum of health science is being fundamentally transformed by the application of basic science informatics in the domains of meanings, ontologies, and natural language processing; these have matured into critical technologies and unquestioned adjuncts of the emerging big-science transformation of biomedicine and healthcare. In this issue of the Journal, two articles illustrate some of these dependencies upon semantic health technologies. The first employs sophisticated dictionary lookup techniques to classify news and specialist articles about disease outbreaks as an adjunct to outbreak detection and severity measurement. The second posits a sophisticated scaling of outbreak severity based not only on disease metrics but also on sociological and governmental reactions in the face of mild to severe epidemics. The article by Freifeld and colleagues1 describes classification and display software that systematically collects data from outbreak-alert distribution lists and general news media in order to classify disease and location. Once classified, the information is visually displayed on a Web-rendered world map with various capabilities for expanding or contracting time windows and locations. The classification engine itself involves a multistage parsing, part of which includes fast dictionary lookups against word-level N-grams using cascading hashes of keywords, locations, and organisms. Additionally, this dictionary has limited ontologic capabilities, such as “containment,” where it will explicitly recognize that the city of Boston is located in the state of Massachusetts. There are significant technical limitations, readily acknowledged by the authors, such as an undue reliance on exact phrase matching. On the other hand, this brute force dictionary lookup methodology has the advantage of generating results that are highly effective, and it illustrates the utility of such Internet surveillance for identifying outbreaks and their locations and for depicting their apparent intensity. The work by Wilson et al.2 focuses on the severity of social disruption for biological outbreaks among humans or animals by examining “the indications and warnings (or markers) of social disruption used in tandem with explicit reports of disease.” The authors' emphasis is on defining stages of outbreak severity, which include (adapted from Table 1): 0) Environmental conditions favorable to an outbreak, 1) Localized biological event, 2) Multi-focal biological event, 3) Severe social and medical infrastructure strain, 4) Social collapse, and P) Preparatory posture. However, in their discussion, the authors explicitly raise the obvious strategy of Internet technologies as possible “harvesting engines” to capture information relevant to the definitional criteria for biological-outbreak severity metrics. The examples they provide on Rift Valley Fever, Venezuelan Equine Encephalitis, and SARS were all humanly annotated and curated. Nevertheless, the prospect of applying algorithmic methods to “harvest” the requisite information from broadly based, multilingual news media on the Internet to 1) identify emerging outbreaks and 2) assign initial outbreak severity or track the escalation of severity over time is the point relevant to this editorial. While these manuscripts differ dramatically in their scope—the first focusing on disambiguating disease and location of outbreaks, the second spanning similar biological characteristics but also including sociological and governmental responses in the face of public health threats—they are similar in their existing or implied dependence on semantic health technologies. As biology and medicine evolve toward a big-science paradigm, analogous to the evolution experienced by physics early in the 20th century or astronomy in the latter half of the century, an obvious question for the authors of these two manuscripts and their underpinning infrastructures is whether any data sharing or interoperability can occur among their described systems—present or future. Efforts to achieve an efficient public health infrastructure must obviously invoke Semantic Web principles to establish shared terminologies, ontologies, and knowledge metadata. It would be intriguing to establish whether the disease ontologies across these systems have a machinable overlap (it is assumed that they have an implicit semantic overlap). Assuming they do not, how can standard ontologies of disease outbreaks evolve? In a similar vein, the systems described in these articles are explicitly or implicitly dependent upon Natural Language Processing (NLP) of news and outbreak reports. To what extent do they share dictionary lookup technologies or more sophisticated language processing capabilities? While these questions are rhetorical in the present circumstance, they increasingly define the research agenda and infrastructure creation needed to achieve the visions captured by these articles. Public health is not alone in its evolution toward big science. Basic biology, anchored on the human genome, has witnessed the emergence of the Gene Ontology as a nascent interlingua to describe biological functions associated with genes and gene products. Similarly, NLP techniques are harvesting the biomedical literature in an effort to build more coherent knowledge atop the computationally fragile collections of quaint text. In clinical medicine, NLP and data normalization form the bedrock of emerging clinical data warehouses, whose purpose explicitly is to improve the quality of care for existing patients and aid the discovery of new knowledge. Transcending these arenas is the Genotype–Phenotype Holy Grail, now manifest by many initiatives, including the NHGRI Genome-wide Association Studies (GWAS) linked with EMR-derived phenotype definitions; this translational science is even more dependent on semantic health technologies for generalizeable success. Among projects that seek to provide unity across this spectrum of genomic to phenotypic characterization, with application in basic science, clinical medicine, and most assuredly public health, is the next generation International Classification of Diseases (ICD) in early development by the World Health Organization (WHO). In my capacity as Chair of the ICD Revision Steering Committee, I can report that, while promising to be built on robust ontologic principles, with linkage to underpinning terminologies, the resulting high-level rubrics or classification spaces of the new ICD may be candidates for the kind of infrastructure that the two articles in today's Journal might effectively leverage—at least with respect to disease. Corresponding work is ongoing with the new SNOMED (now known as the International Health Terminology—IHT), the Gene Ontology, and emergent products from the Open Biomedical Ontologies; all potentially coordinated by the newly established National Center for Biomedical Ontologies based out of Stanford. To make NLP and efficient dictionary lookup techniques practical, a robust thesaurus of standard terminologies and classifications must evolve to include multilingual synonymy as well as phrase-level synonyms within languages. The challenge of creating such an open-source, fundamental linguistic resource for basic biological science, clinical medicine, and public health is slowly being recognized as the sine qua non for harvesting information efficiently and fulfilling the promise of biology and medicine by moving toward genuine big-science integration. Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | Technical Brief: Mayo Clinic NLP System for Patient Smoking Status IdentificationabstractThis article describes our system entry for the 2006 I2B2 contest "Challenges in Natural Language Processing for Clinical Data" for the task of identifying the smoking status of patients. Our system makes the simplifying assumption that patient-level smoking status determination can be achieved by accurately classifying individual sentences from a patient's record. We created our system with reusable text analysis components built on the Unstructured Information Management Architecture and Weka. This reuse of code minimized the development effort related specifically to our smoking status classifier. We report precision, recall, F-score, and 95% exact confidence intervals for each metric. Recasting the classification task for the sentence level and reusing code from other text analysis projects allowed us to quickly build a classification system that performs with a system F-score of 92.64 based on held-out data tests and of 85.57 on the formal evaluation data. Our general medical natural language engine is easily adaptable to a real-world medical informatics application. Some of the limitations as applied to the use-case are negation detection and temporal resolution. Guergana K. Savova, Philip V. Ogren, Patrick H. Duffy, James D. Buntrock, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 5 |
| 2008 | Word sense disambiguation across two domains: Biomedical literature and clinical notes
Guergana K. Savova, Anni Coden, Igor L. Sominsky, Rie Johnson, Philip V. Ogren, Piet C. de Groen, Christopher G. Chute |
J. Biomed. Informatics | 7 |
| 2007 | Measures of semantic similarity and relatedness in the biomedical domain
Ted Pedersen, Serguei V. S. Pakhomov, Siddharth Patwardhan, Christopher G. Chute |
J. Biomed. Informatics | 4 |
| 2006 | An End-to-End Supervised Target-Word Sense Disambiguation System
Mahesh Joshi, Serguei V. S. Pakhomov, Ted Pedersen, Richard Maclin, Christopher G. Chute |
AAAI | 5 |
| 2006 | A Comparative Study of Supervised Learning as Applied to Acronym Expansion in Clinical Reports
Mahesh Joshi, Serguei V. S. Pakhomov, Ted Pedersen, Christopher G. Chute |
AMIA | 4 |
| 2006 | Building and Evaluating Annotated Corpora for Medical NLP Systems
Philip V. Ogren, Guergana K. Savova, James D. Buntrock, Christopher G. Chute |
AMIA | 4 |
| 2006 | A Hybrid Approach to Determining Modification of Clinical Diagnoses
Serguei V. S. Pakhomov, Christopher G. Chute |
AMIA | 2 |
| 2006 | Research Paper: Automating the Assignment of Diagnosis Codes to Patient Encounters Using Example-based and Machine Learning TechniquesabstractOBJECTIVE: Human classification of diagnoses is a labor intensive process that consumes significant resources. Most medical practices use specially trained medical coders to categorize diagnoses for billing and research purposes. METHODS: We have developed an automated coding system designed to assign codes to clinical diagnoses. The system uses the notion of certainty to recommend subsequent processing. Codes with the highest certainty are generated by matching the diagnostic text to frequent examples in a database of 22 million manually coded entries. These code assignments are not subject to subsequent manual review. Codes at a lower certainty level are assigned by matching to previously infrequently coded examples. The least certain codes are generated by a naïve Bayes classifier. The latter two types of codes are subsequently manually reviewed. MEASUREMENTS: Standard information retrieval accuracy measurements of precision, recall and f-measure were used. Micro- and macro-averaged results were computed. RESULTS At least 48% of all EMR problem list entries at the Mayo Clinic can be automatically classified with macro-averaged 98.0% precision, 98.3% recall and an f-score of 98.2%. An additional 34% of the entries are classified with macro-averaged 90.1% precision, 95.6% recall and 93.1% f-score. The remaining 18% of the entries are classified with macro-averaged 58.5%. CONCLUSION: Over two thirds of all diagnoses are coded automatically with high accuracy. The system has been successfully implemented at the Mayo Clinic, which resulted in a reduction of staff engaged in manual coding from thirty-four coders to seven verifiers. Serguei V. S. Pakhomov, James D. Buntrock, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 3 |
| 2005 | LexGrid Editor: Terminology Authoring for the Lexical Grid
Thomas M. Johnson, Harold R. Solbrig, Daniel C. Armbrust, Christopher G. Chute |
AMIA | 4 |
| 2005 | Abbreviation and Acronym Disambiguation in Clinical Discourse
Serguei V. S. Pakhomov, Ted Pedersen, Christopher G. Chute |
AMIA | 3 |
| 2005 | Frame Semantics and the Domain of Functioning, Disability and Health
Guergana K. Savova, Marcelline R. Harris, Serguei V. S. Pakhomov, Christopher G. Chute |
AMIA | 4 |
| 2005 | Representing Lexical Components of Medical Terminologies in OWL
Kaustubh Supekar, Christopher G. Chute, Harold R. Solbrig |
AMIA | 2 |
| 2005 | Domain-specific language models and lexicons for tagging
Anni Coden, Serguei V. S. Pakhomov, Rie Kubota Ando, Patrick H. Duffy, Christopher G. Chute |
J. Biomed. Informatics | 5 |
| 2005 | Prospective recruitment of patients with congestive heart failure using an ad-hoc binary classifier
Serguei V. S. Pakhomov, James D. Buntrock, Christopher G. Chute |
J. Biomed. Informatics | 3 |
| 2003 | Integrating Pharmacokinetics Knowledge into a Drug Ontology As an Extension to Support Pharmacogenomics
Christopher G. Chute, John S. Carter, Mark S. Tuttle, Margaret W. Haber, Steven H. Brown |
AMIA | 1 |
| 2003 | Testing the Generalizability of the ISO Model for Nursing Diagnoses
Marcelline R. Harris, Hyeon-Eui Kim, Lori Rhudy, Guergana K. Savova, Christopher G. Chute |
AMIA | 5 |
| 2003 | A Data-Driven Approach for Extracting "the Most Specific Term" for Ontology Development
Guergana K. Savova, Marcelline R. Harris, Thomas M. Johnson, Serguei V. S. Pakhomov, Christopher G. Chute |
AMIA | 5 |
| 2003 | The Open Terminology Services (OTS) Project: AMIA 2003 Open Source Expo
Harold R. Solbrig, Daniel C. Armbrust, Christopher G. Chute |
AMIA | 3 |
| 2003 | A term extraction tool for expanding content in the domain of functioning, disability, and health: proof of concept
Marcelline R. Harris, Guergana K. Savova, Thomas M. Johnson, Christopher G. Chute |
J. Biomed. Informatics | 4 |
| 2002 | An evaluation of unmediated versus mediated retrieval services
James D. Buntrock, Christopher G. Chute |
AMIA | 2 |
| 2002 | The horizontal and vertical nature of patient phenotype retrieval: new directions for clinical text processing
Christopher G. Chute |
AMIA | 1 |
| 2002 | Maximum entropy modeling for mining patient medication status from free text
Serguei V. S. Pakhomov, Alexander Ruggieri, Christopher G. Chute |
AMIA | 3 |
| 2001 | An assessment of cancer clinical trials vocabulary and IT infrastructure in the U.S
K. Gwaltney, Christopher G. Chute, Douglas Hageman, Warren A. Kibbe, Kathleen A. McCormick, Dianne M. Reeves, Lawrence W. Wright |
AMIA | 2 |
| 2001 | Expression of a domain ontology model in unified modeling language for the World Health Organization International classification of impairment, disability, and handicap, version 2
Alexander Ruggieri, Peter L. Elkin, Harold R. Solbrig, Christopher G. Chute |
AMIA | 4 |
| 2000 | A randomized controlled trial of concept based indexing of Web page content
Peter L. Elkin, Alexander Ruggieri, Larry Bergstrom, Brent A. Bauer, Philip V. Ogren, Christopher G. Chute |
AMIA | 7 |
| 2000 | The content coverage and organizational structure of terminologies: the example of postoperative pain
Marcelline R. Harris, Judith R. Graves, Linda Herrick, Peter L. Elkin, Christopher G. Chute |
AMIA | 5 |
| 2000 | Representation by standard terminologies of health status concepts contained in two health status assessment instruments used in rheumatic disease management
Alexander Ruggieri, Peter L. Elkin, Christopher G. Chute |
AMIA | 3 |
| 2000 | A formal approach to integrating synonyms with a reference terminology
Harold R. Solbrig, Peter L. Elkin, Philip V. Ogren, Christopher G. Chute |
AMIA | 4 |
| 2000 | Viewpoint: Clinical Classification and Terminology: Some History and Current ObservationsabstractThe evolution of health terminology has undergone glacial transition over time, although this pace has quickened recently. After a long history of near neglect, unimaginative structure, and factitious development, health terminologies are in an era of unprecedented importance, sophistication, and collaboration. The major highlights of this history are reviewed, together with important intellectual advances in health terminology development. The inescapable conclusion is that we are amidst a major revolution in the role and capabilities of health terminologies, entering an age of large-scale systems for health concept representation with international implications. Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2000 | Review: Embedded Structures and Representation of Nursing KnowledgeabstractNursing Vocabulary Summit participants were challenged to consider whether reference terminology and information models might be a way to move toward better capture of data in electronic medical records. A requirement of such reference models is fidelity to representations of domain knowledge. This article discusses embedded structures in three different approaches to organizing domain knowledge: scientific reasoning, expertise, and standardized nursing languages. The concept of pressure ulcer is presented as an example of the various ways lexical elements used in relation to a specific concept are organized across systems. Different approaches to structuring information-the clinical information system, minimum data sets, and standardized messaging formats-are similarly discussed. Recommendations include identification of the polyhierarchies and categorical structures required within a reference terminology, systematic evaluations of the extent to which structured information accurately and completely represents domain knowledge, and modifications or extensions to existing multidisciplinary efforts. Marcelline R. Harris, Judith R. Graves, Harold R. Solbrig, Peter L. Elkin, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 5 |
| 1999 | Desiderata for a clinical terminology server
Christopher G. Chute, Peter L. Elkin, David D. Sherertz, Mark S. Tuttle |
AMIA | 1 |
| 1999 | Detailed content and terminological properties of DSM-IV
Dermot Dunne, Christopher G. Chute |
AMIA | 2 |
| 1999 | A randomized double-blind controlled trial of automated term dissection
Peter L. Elkin, Kent R. Bailey, Philip V. Ogren, Brent A. Bauer, Christopher G. Chute |
AMIA | 5 |
| 1999 | Towards Unambiguous Representation of Nursing-Sensitive Concepts
Marcelline R. Harris, Christopher G. Chute |
AMIA | 2 |
| 1999 | A large-scale evaluation of terminology integration characteristics
Furman S. McDonald, Christopher G. Chute, Philip V. Ogren, Dietlind Wahner-Roedler, Peter L. Elkin |
AMIA | 2 |
| 1999 | Barriers to the clinical implementation of compositionality
Lawrence K. McKnight, Peter L. Elkin, Philip V. Ogren, Christopher G. Chute |
AMIA | 4 |
| 1998 | The Copernican era of healthcare terminology: a re-centering of health information systems
Christopher G. Chute |
AMIA | 1 |
| 1998 | A clinical terminology in the post modern era: pragmatic problem list development
Christopher G. Chute, Peter L. Elkin, Susan H. Fenton, Geoffrey E. Atkin |
AMIA | 1 |
| 1998 | A randomized controlled trial of automated term composition
Peter L. Elkin, Kent R. Bailey, Christopher G. Chute |
AMIA | 3 |
| 1998 | A Java-Based Tool for Entry of a Medical Problem List, Which Accesses a Remote Large-Scale Enterprise Vocabulary Server
Peter L. Elkin, Mark S. Tuttle, Kevin Keck, Geoffrey E. Atkin, Christopher G. Chute |
AMIA | 5 |
| 1998 | A Java-Based Interface for Medical Research Project Classification Using Metaphrase
James D. Buntrock, Douglas L. Crowson, Mark S. Tuttle, Peter L. Elkin, Christopher G. Chute |
AMIA | 6 |
| 1998 | Position Paper: A Framework for Comprehensive Health Terminology Systems in the United States: Development Guidelines, Criteria for Selection, and Public Policy ImplicationsabstractHealth care in the United States has become an information-intensive industry, yet electronic health records represent patient data inconsistently for lack of clinical data standards. Classifications that have achieved common acceptance, such as the ICD-9-CM or ICD, aggregate heterogeneous patients into broad categories, which preclude their practical use in decision support, development of refined guidelines, or detailed comparison of patient outcomes or benchmarks. This document proposes a framework for the integration and maturation of clinical terminologies that would have practical applications in patient care, process management, outcome analysis, and decision support. Arising from the two working groups within the standards community--the ANSI (American National Standards Institute) Healthcare Informatics Standards Board Working Group and the Computer-based Patient Records Institute Working Group on Codes and Structures--it outlines policies regarding 1) functional characteristics of practical terminologies, 2) terminology models that can broaden their applications and contribute to their sustainability, 3) maintenance attributes that will enable terminologies to keep pace with rapidly changing health care knowledge and process, and 4) administrative issues that would facilitate their accessibility, adoption, and application to improve the quality and efficiency of American health care. Christopher G. Chute, Simon P. Cohn, James R. Campbell 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 1997 | A clinically derived terminology: qualification to reduction
Christopher G. Chute, Peter L. Elkin |
AMIA | 1 |
| 1997 | Clinical care management and workflow by episodes
P. L. Claus, Paul C. Carpenter, Christopher G. Chute, David N. Mohr, P. S. Gibbons |
AMIA | 3 |
| 1997 | Metaphrase: Achieving Formalized EMR Problem Lists from Informal Input
William G. Cole, David D. Sherertz, Mark S. Tuttle, Kevin Keck, Nels Olson, Christopher G. Chute, Charles Safran |
AMIA | 6 |
| 1997 | Standardized problem list generation, utilizing the Mayo canonical vocabulary embedded within the Unified Medical Language System
Peter L. Elkin, David N. Mohr, Mark S. Tuttle, William G. Cole, Geoffrey E. Atkin, Kevin Keck, Thomas B. Fisk, B. H. Kaihoi, K. E. Lee, Michael C. Higgins, Henri J. Suermondt, Nels Olson, P. L. Claus, Paul C. Carpenter, Christopher G. Chute |
AMIA | 15 |
| 1997 | Supporting Postcoordination in an Electronic Problem List
Kevin Keck, Keith E. Campbell, Christopher G. Chute, Peter L. Elkin, Mark S. Tuttle, William G. Cole |
AMIA | 3 |
| 1997 | Research Paper: Phase II Evaluation of Clinical Coding Schemes: Completeness, Taxonomy, Mapping, Definitions, and ClarityabstractOBJECTIVE: To compare three potential sources of controlled clinical terminology (READ codes version 3.1, SNOMED International, and Unified Medical Language System (UMLS) version 1.6) relative to attributes of completeness, clinical taxonomy, administrative mapping, term definitions and clarity (duplicate coding rate). METHODS: The authors assembled 1929 source concept records from a variety of clinical information taken from four medical centers across the United States. The source data included medical as well as ample nursing terminology. The source records were coded in each scheme by an investigator and checked by the coding scheme owner. The codings were then scored by an independent panel of clinicians for acceptability. Codes were checked for definitions provided with the scheme. Codes for a random sample of source records were analyzed by an investigator for "parent" and "child" codes within the scheme. Parent and child pairs were scored by an independent panel of medical informatics specialists for clinical acceptability. Administrative and billing code mapping from the published scheme were reviewed for all coded records and analyzed by independent reviewers for accuracy. The investigator for each scheme exhaustively searched a sample of coded records for duplications. RESULTS: SNOMED was judged to be significantly more complete in coding the source material than the other schemes (SNOMED* 70%; READ 57%; UMLS 50%; *p < .00001). SNOMED also had a richer clinical taxonomy judged by the number of acceptable first-degree relatives per coded concept (SNOMED* 4.56, UMLS 3.17; READ 2.14, *p < .005). Only the UMLS provided any definitions; these were found for 49% of records which had a coding assignment. READ and UMLS had better administrative mappings (composite score: READ* 40.6%; UMLS* 36.1%; SNOMED 20.7%, *p < .00001), and SNOMED had substantially more duplications of coding assignments (duplication rate: READ 0%; UMLS 4.2%; SNOMED* 13.9%, *p < .004) associated with a loss of clarity. CONCLUSION: No major terminology source can lay claim to being the ideal resource for a computer-based patient record. However, based upon this analysis of releases for April 1995, SNOMED International is considerably more complete, has a compositional nature and a richer taxonomy. Is suffers from less clarity, resulting from a lack of syntax and evolutionary changes in its coding scheme. READ has greater clarity and better mapping to administrative schemes (ICD-10 and OPCS-4), is rapidly changing and is less complete. UMLS is a rich lexical resource, with mappings to many source vocabularies. It provides definitions for many of its terms. However, due to the varying granularities and purposes of its source schemes, it has limitations for representation of clinical concepts within a computer-based patient record. James R. Campbell 0001, Paul C. Carpenter, Charles Sneiderman, Simon P. Cohn, Christopher G. Chute, Judith J. Warren |
J. Am. Medical Informatics Assoc. | 5 |
| 1996 | Research Paper: The Content Coverage of Clinical classificationsabstractBACKGROUND AND OBJECTIVE: Patient conditions and events are the core of patient record content. Computer-based records will require standard vocabularies to represent these data consistently, thereby facilitating clinical decision support, research, and efficient care delivery. To address whether existing major coding systems can serve this function, the authors evaluated major clinical classifications for their content coverage. METHODS: Clinical text from four medical centers was sampled from inpatient and outpatient settings. The resultant corpus of 14,247 words was parsed into 3,061 distinct concepts. These concepts were grouped into Diagnoses, Modifiers, Findings, Treatments and Procedures, and Other. Each concept was coded into ICD-9-CM, ICD-10, CPT, SNOMED III, Read V2, UMLS 1.3, and NANDA; a secondary reviewer ensured consistency. While coding, the information was scored: 0 = no match, 1 = fair match, 2 = complete match. RESULTS: ICD-9-CM had an overall mean score of 0.77 out of 2; its highest subscore was 1.61 for Diagnoses. ICD-10 scored 1.60 for Diagnoses, and 0.62 overall. The overall score of ICD-9-CM augmented by CPT was not materially improved at 0.82. The SNOMED International system demonstrated the highest score in every category, including Diagnoses (1.90), and had an overall score of 1.74. CONCLUSION: No classification captured all concepts, although SNOMED did notably the most complete job. The systems in major use in the United States, ICD-9-CM and CPT, fail to capture substantial clinical content. ICD-10 does not perform better than ICD-9-CM. The major clinical classifications in use today incompletely cover the clinical content of patient records; thus analytic conclusions that depend on these systems may be suspect. Christopher G. Chute, Simon P. Cohn, Keith E. Campbell, Diane E. Oliver, James R. Campbell 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 1994 | An Example-Based Mapping Method for Text Categorization and RetrievalabstractA unified model for text categorization and text retrieval is introduced. We use a training set of manually categorized documents to learn word-category associations, and use these associations to predict the categories of arbitrary documents. Similarly, we use a training set of queries and their related documents to obtain empirical associations between query words and indexing terms of documents, and use these associations to predict the related documents of arbitrary queries. A Linear Least Squares Fit (LLSF) technique is employed to estimate the likelihood of these associations. Document collections from the MEDLINE database and Mayo patient records are used for studies on the effectiveness of our approach, and on how much the effectiveness depends on the choices of training data, indexing language, word-weighting scheme, and morphological canonicalization. Alternative methods are also tested on these data collections for comparison. It is evident that the LLSF approach uses the relevance information effectively within human decisions of categorization and retrieval, and achieves a semantic mapping of free texts to their representations in an indexing language. Such a semantic mapping lead to a significant improvement in categorization and retrieval, compared to alternative approaches. Yiming Yang 0002, Christopher G. Chute |
ACM Trans. Inf. Syst. | 2 |
| 1993 | An Application of Least Squares Fit Mapping to Text Information RetrievalabstractThis paper describes a unique example-based mapping method for document retrieval. We discovered that the knowledge about relevance among queries and documents can be used to obtain empirical connections between query terms and the canonical concepts which are used for indexing the content of documents. These connections do not depend on whether there are shared terms among the queries and documents; therefore, they are especially effective for a mapping from queries to the documents where the concepts are relevant but the terms used by article authors happen to be different from the terms of database users. We employ a Linear Least Squares Fit (LLSF) technique to compute such connections from a collection of queries and documents where the relevance is assigned by humans, and then use these connections in the retrieval of documents where the relevance is unknown. We tested this method on both retrieval and indexing with a set of MEDLINE documents which has been used by other information retrieval systems for evaluations. The effectiveness of the LLSF mapping and the significant improvement over alternative approaches was evident in the tests. Yiming Yang 0002, Christopher G. Chute |
SIGIR | 2 |
| 1992 | A Linear Least Squares Fit Mapping Method For Information Retrieval From Natural Language Texts
Yiming Yang 0002, Christopher G. Chute |
COLING | 2 |