Karen Eilbeck

dblp:64/1762 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-0831-6427ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 3 first-author · 4 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 GARDE-Chat: a scalable, open-source platform for building and deploying health chatbots
abstract
BACKGROUND: Chatbots are increasingly used to deliver health education, patient engagement, and access to healthcare services. GARDE-Chat is an open-source platform designed to facilitate the development, deployment, and dissemination of chatbot-based digital health interventions across different domains and settings. MATERIALS AND METHODS: GARDE-Chat was developed through an iterative process informed by real-world use cases to guide prioritization of key features. The tool was developed as an open-source platform to promote collaboration, broad dissemination, and impact across research and clinical domains. RESULTS: GARDE-Chat's main features include (1) a visual authoring interface that allows non-programmers to design chatbots; (2) support for scripted, large language model (LLM)-based and hybrid chatbots; (3) capacity to share chatbots with researchers and institutions; (4) integration with external applications and data sources such as electronic health records and REDCap; (5) delivery via web browsers or text messaging; and (6) detailed audit log supporting analyses of chatbot user interactions. Since its first release in July 2022, GARDE-Chat has supported the development of chatbot-based interventions tested in multiple studies, including large pragmatic clinical trials addressing topics such as genetic testing, COVID-19 testing, tobacco cessation, and cancer screening. DISCUSSION: Ongoing challenges include the effort required for developing chatbot scripts, ensuring safe use of LLMs, and integrating with clinical systems. CONCLUSION: GARDE-Chat is a generalizable platform for creating, implementing, and disseminating scalable chatbot-based population health interventions. It has been validated in several studies, and it is available to researchers and healthcare systems through an open-source mechanism.
Guilherme Del Fiol, Emerson P. Borsato, Richard L. Bradshaw, Jiantao Bian, Alana Woodbury, Courtney Gauchel, Karen Eilbeck, Whitney Maxwell, Kelsey Ellis, Anne C. Madeo, Chelsey R. Schlechter, Polina V. Kukhareva, Caitlin G. Allen, Michael Kean, Elena B. Elkin, Ravi Sharaf, Muhammad D. Ahsan, Melissa Frey, Lauren Davis-Rivera, Wendy Kohlmann, David W. Wetter, Kimberly A. Kaphingst, Kensaku Kawamoto
J. Am. Medical Informatics Assoc.7
2024 The Business Process Management for Healthcare (BPM+ Health) Consortium: motivation, methodology, and deliverables for enabling clinical knowledge interoperability (CKI)
abstract
OBJECTIVES: To enhance the Business Process Management (BPM)+ Healthcare language portfolio by incorporating knowledge types not previously covered and to improve the overall effectiveness and expressiveness of the suite to improve Clinical Knowledge Interoperability. METHODS: We used the BPM+ Health and Object Management Group (OMG) standards development methodology to develop new languages, following a gap analysis between existing BPM+ Health languages and clinical practice guideline knowledge types. Proposal requests were developed based on these requirements, and submission teams were formed to respond to them. The resulting proposals were submitted to OMG for ratification. RESULTS: The BPM+ Health family of languages, which initially consisted of the Business Process Model and Notation, Decision Model and Notation, and Case Model and Notation, was expanded by adding 5 new language standards through the OMG. These include Pedigree and Provenance Model and Notation for expressing epistemic knowledge, Knowledge Package Model and Notation for supporting packaging knowledge, Shared Data Model and Notation for expressing ontic knowledge, Party Model and Notation for representing entities and organizations, and Specification Common Elements, a language providing a standard abstract and reusable library that underpins the 4 new languages. DISCUSSION AND CONCLUSION: In this effort, we adopted a strategy of separation of concerns to promote a portfolio of domain-agnostic, independent, but integrated domain-specific languages for authoring medical knowledge. This strategy is a practical and effective approach to expressing complex medical knowledge. These new domain-specific languages offer various knowledge-type options for clinical knowledge authors to choose from without potentially adding unnecessary overhead or complexity.
Robert F. Lario, Richard Soley, Stephen White, John Butler, Guilherme Del Fiol, Karen Eilbeck, Stanley M. Huff, Kensaku Kawamoto
J. Am. Medical Informatics Assoc.6
2023 A method for structuring complex clinical knowledge and its representational formalisms to support composite knowledge interoperability in healthcare
abstract
INTRODUCTION: The use and interoperability of clinical knowledge starts with the quality of the formalism utilized to express medical expertise. However, a crucial challenge is that existing formalisms are often suboptimal, lacking the fidelity to represent complex knowledge thoroughly and concisely. Often this leads to difficulties when seeking to unambiguously capture, share, and implement the knowledge for care improvement in clinical information systems used by providers and patients. OBJECTIVES: To provide a systematic method to address some of the complexities of knowledge composition and interoperability related to standards-based representational formalisms of medical knowledge. METHODS: Several cross-industry (Healthcare, Linguistics, System Engineering, Standards Development, and Knowledge Engineering) frameworks were synthesized into a proposed reference knowledge framework. The framework utilizes IEEE 42010, the MetaObject Facility, the Semantic Triangle, an Ontology Framework, and the Domain and Comprehensibility Appropriateness criteria. The steps taken were: 1) identify foundational cross-industry frameworks, 2) select architecture description method, 3) define life cycle viewpoints, 4) define representation and knowledge viewpoints, 5) define relationships between neighboring viewpoints, and 6) establish characteristic definitions of the relationships between components. System engineering principles applied included separation of concerns, cohesion, and loose coupling. RESULTS: A "Multilayer Metamodel for Representation and Knowledge" (M*R/K) reference framework was defined. It provides a standard vocabulary for organizing and articulating medical knowledge curation perspectives, concepts, and relationships across the artifacts created during the life cycle of language creation, authoring medical knowledge, and knowledge implementation in clinical information systems such as electronic health records (EHR). CONCLUSION: M*R/K provides a systematic means to address some of the complexities of knowledge composition and interoperability related to medical knowledge representations used in diverse standards. The framework may be used to guide the development, assessment, and coordinated use of knowledge representation formalisms. M*R/K could promote the alignment and aggregated use of distinct domain-specific languages in composite knowledge artifacts such as clinical practice guidelines (CPGs).
Robert F. Lario, Kensaku Kawamoto, Davide Sottara, Karen Eilbeck, Stanley M. Huff, Guilherme Del Fiol, Richard Soley, Blackford Middleton
J. Biomed. Informatics4
2022 LocalVar: a local variant collection manager to asynchronously detect synonyms, HGVS expression changes, and variant interpretation changes from ClinVar
Michael T. Watkins, Wendy Kohlmann, Therese Berry, Neetha Sama, Cathryn Koptiuch, Shawn Rynearson, Karen Eilbeck
AMIA7
2020 Utilization of BPM+ Health for the Representation of Clinical Knowledge: A Framework for the Expression and Assessment of Clinical Practice Guidelines (CPG) Utilizing Existing and Emerging Object Management Group (OMG) Standards
Robert F. Lario, Steve Hasely, Stephen White, Karen Eilbeck, Richard Soley, Stanley M. Huff, Kensaku Kawamoto
AMIA4
2020 Unification of miRNA and isomiR research: the mirGFF3 format and the mirtop API
abstract
MOTIVATION: MicroRNAs (miRNAs) are small RNA molecules (∼22 nucleotide long) involved in post-transcriptional gene regulation. Advances in high-throughput sequencing technologies led to the discovery of isomiRs, which are miRNA sequence variants. While many miRNA-seq analysis tools exist, the diversity of output formats hinders accurate comparisons between tools and precludes data sharing and the development of common downstream analysis methods. RESULTS: To overcome this situation, we present here a community-based project, miRNA Transcriptomic Open Project (miRTOP) working towards the optimization of miRNA analyses. The aim of miRTOP is to promote the development of downstream isomiR analysis tools that are compatible with existing detection and quantification tools. Based on the existing GFF3 format, we first created a new standard format, mirGFF3, for the output of miRNA/isomiR detection and quantification results from small RNA-seq data. Additionally, we developed a command line Python tool, mirtop, to create and manage the mirGFF3 format. Currently, mirtop can convert into mirGFF3 the outputs of commonly used pipelines, such as seqbuster, isomiR-SEA, sRNAbench, Prost! as well as BAM files. Some tools have also incorporated the mirGFF3 format directly into their code, such as, miRge2.0, IsoMIRmap and OptimiR. Its open architecture enables any tool or pipeline to output or convert results into mirGFF3. Collectively, this isomiR categorization system, along with the accompanying mirGFF3 and mirtop API, provide a comprehensive solution for the standardization of miRNA and isomiR annotation, enabling data sharing, reporting, comparative analyses and benchmarking, while promoting the development of common miRNA methods focusing on downstream steps of miRNA detection, annotation and quantification. AVAILABILITY AND IMPLEMENTATION: https://github.com/miRTop/mirGFF3/ and https://github.com/miRTop/mirtop. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thomas Desvignes, Phillipe Loher, Karen Eilbeck, Jeffery Ma, Gianvito Urgese, Bastian Fromm, Jason Sydes, Ernesto Aparicio-Puerta, Víctor Barrera, Roderic Espín, Florian Thibord, Xavier Bofill-De Ros, Eric Londin, Aristeidis G. Telonis, Elisa Ficarra, Marc R. Friedländer, John H. Postlethwait, Isidore Rigoutsos, Michael Hackenberg, Ioannis S. Vlachos, Marc K. Halushka, Lorena Pantano
Bioinform.3
2019 Continuous Predictive Modelling for Sepsis, Organ Failure, and In-hospital Mortality in the Intensive Care Unit
Amber Kiser, Devin Horton, Karen Eilbeck, Samir E. AbdelRahman
AMIA3
2019 Implementing the VMC specification to reduce ambiguity in genomic variant representation
Michael T. Watkins, Shawn Rynearson, Alex Henrie, Karen Eilbeck
AMIA4
2018 The VAAST Variant Prioritizer (VVP): ultrafast, easy to use whole genome variant prioritization tool
abstract
BACKGROUND: Prioritization of sequence variants for diagnosis and discovery of Mendelian diseases is challenging, especially in large collections of whole genome sequences (WGS). Fast, scalable solutions are needed for discovery research, for clinical applications, and for curation of massive public variant repositories such as dbSNP and gnomAD. In response, we have developed VVP, the VAAST Variant Prioritizer. VVP is ultrafast, scales to even the largest variant repositories and genome collections, and its outputs are designed to simplify clinical interpretation of variants of uncertain significance. RESULTS: We show that scoring the entire contents of dbSNP (> 155 million variants) requires only 95 min using a machine with 4 cpus and 16 GB of RAM, and that a 60X WGS can be processed in less than 5 min. We also demonstrate that VVP can score variants anywhere in the genome, regardless of type, effect, or location. It does so by integrating sequence conservation, the type of sequence change, allele frequencies, variant burden, and zygosity. Finally, we also show that VVP scores are consistently accurate, and easily interpreted, traits not shared by many commonly used tools such as SIFT and CADD. CONCLUSIONS: VVP provides rapid and scalable means to prioritize any sequence variant, anywhere in the genome, and its scores are designed to facilitate variant interpretation using ACMG and NHS guidelines. These traits make it well suited for operation on very large collections of WGS sequences.
Steven Flygare, Edgar Javier Hernandez, Lon Phan, Barry Moore, Anthony P. Fejes, Karen Eilbeck, Chad D. Huff, Lynn Jorde, Martin G. Reese, Mark Yandell
BMC Bioinform.8
2016 The genomic CDS sandbox: An assessment among domain experts
Ayesha Aziz, Kensaku Kawamoto, Karen Eilbeck, Marc S. Williams, Robert R. Freimuth, Mark A. Hoffman, Luke V. Rasmussen, Casey Overby Taylor, Brian H. Shirts, James M. Hoffman, Brandon M. Welch
J. Biomed. Informatics3
2015 A domain ontology for the Non-Coding RNA field
abstract
Identification of non-coding RNAs (ncRNAs) has been significantly enhanced due to the rapid advancement in sequencing technologies. On the other hand, semantic annotation of ncRNA data lag behind their identification, and there is a great need to effectively integrate discovery from relevant communities. To this end, the Non-Coding RNA Ontology (NCRO) is being developed to provide a precisely defined ncRNA controlled vocabulary, which can fill a specific and highly needed niche in unification of ncRNA biology.
Jingshan Huang, Karen Eilbeck, Judith A. Blake, Dejing Dou, Darren A. Natale, Alan Ruttenberg, Barry Smith 0001, Michael T. Zimmermann, Guoqian Jiang, Bin Wu 0008, Yongqun He, Shaojie Zhang 0001, Xiaowei Wang 0006, Zixing Liu
BIBM2
2015 A semantic approach for knowledge capture of MIcroRNA-Target gene interactions
abstract
Research has indicated that microRNAs (miRNAs), a special class of non-coding RNAs (ncRNAs), can perform important roles in different biological and pathological processes. miRNAs' functions are realized by regulating their respective target genes (targets). It is thus critical to identify and analyze miRNA-target interactions for a better understanding and delineation of miRNAs' functions. However, conventional knowledge discovery and acquisition methods have many limitations. Fortunately, semantic technologies that are based on domain ontologies can render great assistance in this regard. In our previous investigations, we developed a miRNA domain-specific application ontology, Ontology for MIcroRNA Target (OMIT), to provide the community with common data elements and data exchange standards in the miRNA research. This paper describes (1) our continuing efforts in the OMIT ontology development and (2) the application of the OMIT to enable a semantic approach for knowledge capture of miRNA-target interactions.
Jingshan Huang, Fernando Gutierrez, Dejing Dou, Judith A. Blake, Karen Eilbeck, Darren A. Natale, Barry Smith 0001, Xiaowei Wang 0006, Zixing Liu, Alan Ruttenberg
BIBM5
2015 Birth of identity: understanding changes to birth certificates and their value for identity resolution
abstract
INTRODUCTION: Identity information is often used to link records within or among information systems in public health and clinical settings. The quality and stability of birth certificate identifiers impacts both the success of linkage efforts and the value of birth certificate registries for identity resolution. OBJECTIVE: Our objectives were to describe: (1) the frequency and cause of changes to birth certificate identifiers as children age, and (2) the frequency of events (ie, adoptions, paternities, amendments) that may trigger changes and their impact on names. METHODS: We obtained two de-identified datasets from the Utah birth certificate registry: (1) change history from 2000 to 2012, and (2) occurrences for adoptions, paternities, and amendments among births in 1987 and 2000. We conducted cohort analyses for births in 1987 and 2000, examining the number, reason, and extent of changes over time. We conducted cross-sectional analyses to assess the patterns of changes between 2000 and 2012. RESULTS: In a cohort of 48 350 individuals born in 2000 in Utah, 3164 (6.5%) experienced a change in identifiers prior to their 13th birthday, with most changes occurring before 2 years of age. Cross-sectional analysis showed that identifiers are stable for individuals over 5 years of age, but patterns of changes fluctuate considerably over time, potentially due to policy and social factors. CONCLUSIONS: Identities represented in birth certificates change over time. Specific events that cause changes to birth certificates also fluctuate over time. Understanding these changes can help in the development of automated strategies to improve identity resolution.
Jeffrey Duncan, Scott P. Narus, Stephen W. Clyde, Karen Eilbeck, Sidney N. Thornton, Catherine J. Staes
J. Am. Medical Informatics Assoc.4
2014 Evaluation of need for ontologies to manage domain content for the Reportable Conditions Knowledge Management System
Karen Eilbeck, Julie Lipstein, Sunanda R. McGarvey, Catherine J. Staes
AMIA1
2014 Clinical Decision Support for Whole Genome Sequence Information Leveraging a Service-Oriented Architecture: a Prototype
Brandon M. Welch, Salvador Loya, Karen Eilbeck, Kensaku Kawamoto
AMIA3
2014 Technical desiderata for the integration of genomic data with clinical decision support
Brandon M. Welch, Karen Eilbeck, Guilherme Del Fiol, Laurence J. Meyer, Kensaku Kawamoto
J. Biomed. Informatics2
2013 Using KaOS Ontologies to Model Policy Requirements for a Statewide Master Person Index
Jeffrey Duncan, Karen Eilbeck, Catherine J. Staes, Scott P. Narus, Stephen W. Clyde
AMIA2
2011 Evolution of the Sequence Ontology terms and relationships
Chris Mungall, Colin R. Batchelor, Karen Eilbeck
J. Biomed. Informatics3
2009 Quantitative measures for the management and comparison of annotated genomes
abstract
BACKGROUND: The ever-increasing number of sequenced and annotated genomes has made management of their annotations a significant undertaking, especially for large eukaryotic genomes containing many thousands of genes. Typically, changes in gene and transcript numbers are used to summarize changes from release to release, but these measures say nothing about changes to individual annotations, nor do they provide any means to identify annotations in need of manual review. RESULTS: In response, we have developed a suite of quantitative measures to better characterize changes to a genome's annotations between releases, and to prioritize problematic annotations for manual review. We have applied these measures to the annotations of five eukaryotic genomes over multiple releases -- H. sapiens, M. musculus, D. melanogaster, A. gambiae, and C. elegans. CONCLUSION: Our results provide the first detailed, historical overview of how these genomes' annotations have changed over the years, and demonstrate the usefulness of these measures for genome annotation management.
Karen Eilbeck, Barry Moore, Carson Holt, Mark Yandell
BMC Bioinform.1
2008 The Protein Feature Ontology: a tool for the unification of protein feature annotations
abstract
MOTIVATION: The advent of sequencing and structural genomics projects has provided a dramatic boost in the number of uncharacterized protein structures and sequences. Consequently, many computational tools have been developed to help elucidate protein function. However, such services are spread throughout the world, often with standalone web pages. Integration of these methods is needed and so far this has not been possible as there was no common vocabulary available that could be used as a standard language. RESULTS: The Protein Feature Ontology has been developed to provide a structured controlled vocabulary for features on a protein sequence or structure and comprises approximately 100 positional terms, now integrated into the Sequence Ontology (SO) and 40 non-positional terms which describe features relating to the whole-protein sequence. In addition, post-translational modifications are described by using a pre-existing ontology, the Protein Modification Ontology (MOD). This ontology is being used to integrate over 150 distinct annotations provided by the BioSapiens Network of Excellence, a consortium comprising 19 partner sites in Europe. AVAILABILITY: The Protein Feature Ontology can be browsed by accessing the ontology lookup service at the European Bioinformatics Institute (http://www.ebi.ac.uk/ontology-lookup/browse.do?ontName=BS).
Gabrielle A. Reeves, Karen Eilbeck, Michele Magrane, Claire O'Donovan, Luisa Montecchi-Palazzi, Midori A. Harris, Sandra E. Orchard, Rafael C. Jiménez, Andreas Prlic, Tim J. P. Hubbard, Henning Hermjakob, Janet M. Thornton
Bioinform.2
2001 GIMS - A Data Warehouse for Storage and Analysis of Genome Sequence and Functional Data
abstract
Effective analysis of genome sequences and associated functional data requires access to many different kinds of biological information. For example, when analysing gene expression data, it may be useful to have access to the sequences upstream of the genes, or to the cellular location of their protein products. Such information is currently stored in different formats at different sites in a way that does not readily allow integrated analyses. The Genome Information Management System (GIMS) is an object database that integrates genome sequence data with functional data on the transcriptome and on protein-protein interactions in a single data warehouse. We have used GIMS to store the Saccharomyces cerevisiae (yeast) genome and to demonstrate how the integrated storage of diverse kinds of genomic data can be beneficial for analysing data using context-rich queries and analyses. GIMS allows data to be stored in a way that reflects the underlying mechanisms in the organism, and permits complex questions to be asked of the data. This paper provides an overview of the GIMS system and describes some analyses that illustrate its use for analysing functional data sets for S. cerevisiae.
Mike Cornell, Norman W. Paton, Shengli Wu 0001, Carole A. Goble, Crispin J. Miller, Paul Kirby, Karen Eilbeck, Andy Brass, Andrew Hayes, Stephen G. Oliver
BIBE7
2000 Conceptual modelling of genomic information
abstract
Abstract Motivation: Genome sequencing projects are making available complete records of the genetic make-up of organisms. These core data sets are themselves complex, and present challenges to those who seek to store, analyse and present the information. However, in addition to the sequence data, high throughput experiments are making available distinctive new data sets on protein interactions, the phenotypic consequences of gene deletions, and on the transcriptome, proteome, and metabolome. The effective description and management of such data is of considerable importance to bioinformatics in the post-genomic era. The provision of clear and intuitive models of complex information is surprisingly challenging, and this paper presents conceptual models for a range of important emerging information resources in bioinformatics. It is hoped that these can be of benefit to bioinformaticians as they attempt to integrate genetic and phenotypic data with that from genomic sequences, in order to both assign gene functions and elucidate the different pathways of gene action and interaction. Results: This paper presents a collection of conceptual (i.e. implementation-independent) data models for genomic data. These conceptual models are amenable to (more or less direct) implementation on different computing platforms. Availability: Most of the information models presented here have been implemented by the authors using an object database. The implementation of a public interface to this database is in progress. We hope to have a public release in the autumn of 2000, available from http://img.cs.man.ac.uk/gims. Contact: [email protected]
Norman W. Paton, Shakeel A. Khan, Andrew Hayes, Fouzia Moussouni-Marzolf, Andy Brass, Karen Eilbeck, Carole A. Goble, Simon J. Hubbard, Stephen G. Oliver
Bioinform.6
1999 Straight-Line Drawings of Protein Interactions
Wojciech Basalaj, Karen Eilbeck
GD2
1999 INTERACT: An Object Oriented Protein-Protein Interaction Database
Karen Eilbeck, Andy Brass, Norman W. Paton, Charlie Hodgman
ISMB1