EDBT 2026 Demo / reviewers in the wild / expert
Yongqun He
dblp:71/5419
· DBLP profile ↗
37ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-9189-9661ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OntoGrid: Supporting Analysis of Complex Associations in Biomedical Ontologies
Michael Wybrow, Yuan-Fang Li, Tobias Czauderna, Yongqun He |
PacificVis | 5 |
| 2026 | AcuKG: a comprehensive knowledge graph for medical acupunctureabstractBACKGROUND: Acupuncture, a key modality in traditional Chinese medicine, is gaining global recognition as a complementary therapy and a subject of increasing scientific interest. However, fragmented and unstructured acupuncture knowledge spread across diverse sources poses challenges for semantic retrieval, reasoning, and in-depth analysis. To address this gap, we developed AcuKG, a comprehensive knowledge graph that systematically organizes acupuncture-related knowledge to support sharing, discovery, and artificial intelligence-driven innovation in the field. METHODS: AcuKG integrates data from multiple sources, including online resources, guidelines, PubMed literature, ClinicalTrials.gov, and multiple ontologies (SNOMED CT, UBERON, and MeSH). We employed entity recognition, relation extraction, and ontology mapping to establish AcuKG, with human-in-the-loop to ensure data quality. Two cases evaluated AcuKG's usability: (1) how AcuKG advances acupuncture research for obesity and (2) how AcuKG enhances large language model (LLM) application on acupuncture question-answering. RESULTS: AcuKG comprises 1839 entities and 11 527 relations, mapped to 1836 standard concepts in 3 ontologies. Two use cases demonstrated AcuKG's effectiveness and potential in advancing acupuncture research and supporting LLM applications. In the obesity use case, AcuKG identified highly relevant acupoints (eg, ST25, ST36) and uncovered novel research insights based on evidence from clinical trials and literature. When applied to LLMs in answering acupuncture-related questions, integrating AcuKG with GPT-4o and LLaMA 3 significantly improved accuracy (GPT-4o: 46% → 54%, P = .03; LLaMA 3: 17% → 28%, P = .01). CONCLUSION: AcuKG is an open dataset that provides a structured and computational framework for acupuncture applications, bridging traditional practices with acupuncture research and cutting-edge LLM technologies. Xueqing Peng, Su-Yuan Peng, Jianfu Li, Donghong Pei, Fang Li 0011, Yongqun He, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)abstractDespite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts. Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 16 |
| 2022 | COVID-19 vaccine design using reverse and structural vaccinology, ontology-based literature mining and machine learningabstractRational vaccine design, especially vaccine antigen identification and optimization, is critical to successful and efficient vaccine development against various infectious diseases including coronavirus disease 2019 (COVID-19). In general, computational vaccine design includes three major stages: (i) identification and annotation of experimentally verified gold standard protective antigens through literature mining, (ii) rational vaccine design using reverse vaccinology (RV) and structural vaccinology (SV) and (iii) post-licensure vaccine success and adverse event surveillance and its usage for vaccine design. Protegen is a database of experimentally verified protective antigens, which can be used as gold standard data for rational vaccine design. RV predicts protective antigen targets primarily from genome sequence analysis. SV refines antigens through structural engineering. Recently, RV and SV approaches, with the support of various machine learning methods, have been applied to COVID-19 vaccine design. The analysis of post-licensure vaccine adverse event report data also provides valuable results in terms of vaccine safety and how vaccines should be used or paused. Ontology standardizes and incorporates heterogeneous data and knowledge in a human- and computer-interpretable manner, further supporting machine learning and vaccine design. Future directions on rational vaccine design are discussed. Anthony Huffman, Edison Ong, Junguk Hur, Adonis D'mello, Hervé Tettelin, Yongqun He |
Briefings Bioinform. | 6 |
| 2021 | Development of the International Classification of Diseases Ontology (ICDO) and its application for COVID-19 diagnostic data analysisabstractBACKGROUND: The 10th and 9th revisions of the International Statistical Classification of Diseases and Related Health Problems (ICD10 and ICD9) have been adopted worldwide as a well-recognized norm to share codes for diseases, signs and symptoms, abnormal findings, etc. The international Consortium for Clinical Characterization of COVID-19 by EHR (4CE) website stores diagnosis COVID-19 disease data using ICD10 and ICD9 codes. However, the ICD systems are difficult to decode due to their many shortcomings, which can be addressed using ontology. METHODS: An ICD ontology (ICDO) was developed to logically and scientifically represent ICD terms and their relations among different ICD terms. ICDO is also aligned with the Basic Formal Ontology (BFO) and reuses terms from existing ontologies. As a use case, the ICD10 and ICD9 diagnosis data from the 4CE website were extracted, mapped to ICDO, and analyzed using ICDO. RESULTS: We have developed the ICDO to ontologize the ICD terms and relations. Different from existing disease ontologies, all ICD diseases in ICDO are defined as disease processes to describe their occurrence with other properties. The ICDO decomposes each disease term into different components, including anatomic entities, process profiles, etiological causes, output phenotype, etc. Over 900 ICD terms have been represented in ICDO. Many ICDO terms are presented in both English and Chinese. The ICD10/ICD9-based diagnosis data of over 27,000 COVID-19 patients from 5 countries were extracted from the 4CE. A total of 917 COVID-19-related disease codes, each of which were associated with 1 or more cases in the 4CE dataset, were mapped to ICDO and further analyzed using the ICDO logical annotations. Our study showed that COVID-19 targeted multiple systems and organs such as the lung, heart, and kidney. Different acute and chronic kidney phenotypes were identified. Some kidney diseases appeared to result from other diseases, such as diabetes. Some of the findings could only be easily found using ICDO instead of ICD9/10. CONCLUSIONS: ICDO was developed to ontologize ICD10/10 codes and applied to study COVID-19 patient diagnosis data. Our findings showed that ICDO provides a semantic platform for more accurate detection of disease profiles. Ling Wan, Justin Song, Virginia He, Jennifer Roman, Grace Whah, Su-Yuan Peng, Luxia Zhang, Yongqun He |
BMC Bioinform. | 8 |
| 2021 | Visual comprehension and orientation into the COVID-19 CIDO ontology
Yehoshua Perl, Yongqun He, Christopher Ochs, James Geller, Hao Liu 0025, Vipina Kuttichi Keloth |
J. Biomed. Informatics | 3 |
| 2020 | Vaxign-ML: supervised machine learning reverse vaccinology model for improved prediction of bacterial protective antigensabstractMOTIVATION: Reverse vaccinology (RV) is a milestone in rational vaccine design, and machine learning (ML) has been applied to enhance the accuracy of RV prediction. However, ML-based RV still faces challenges in prediction accuracy and program accessibility. RESULTS: This study presents Vaxign-ML, a supervised ML classification to predict bacterial protective antigens (BPAgs). To identify the best ML method with optimized conditions, five ML methods were tested with biological and physiochemical features extracted from well-defined training data. Nested 5-fold cross-validation and leave-one-pathogen-out validation were used to ensure unbiased performance assessment and the capability to predict vaccine candidates against a new emerging pathogen. The best performing model (eXtreme Gradient Boosting) was compared to three publicly available programs (Vaxign, VaxiJen, and Antigenic), one SVM-based method, and one epitope-based method using a high-quality benchmark dataset. Vaxign-ML showed superior performance in predicting BPAgs. Vaxign-ML is hosted in a publicly accessible web server and a standalone version is also available. AVAILABILITY AND IMPLEMENTATION: Vaxign-ML website at http://www.violinet.org/vaxign/vaxign-ml, Docker standalone Vaxign-ML available at https://hub.docker.com/r/e4ong1031/vaxign-ml and source code is available at https://github.com/VIOLINet/Vaxign-ML-docker. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Edison Ong, Haihe Wang, Mei U. Wong, Meenakshi Seetharaman, Ninotchka Valdez, Yongqun He |
Bioinform. | 6 |
| 2020 | Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical eventsabstractOBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making. Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | OntoPlot: A Novel Visualisation for Non-hierarchical Associations in Large OntologiesabstractOntologies are formal representations of concepts and complex relationships among them. They have been widely used to capture comprehensive domain knowledge in areas such as biology and medicine, where large and complex ontologies can contain hundreds of thousands of concepts. Especially due to the large size of ontologies, visualisation is useful for authoring, exploring and understanding their underlying data. Existing ontology visualisation tools generally focus on the hierarchical structure, giving much less emphasis to non-hierarchical associations. In this paper we present OntoPlot, a novel visualisation specifically designed to facilitate the exploration of all concept associations whilst still showing an ontology's large hierarchical structure. This hybrid visualisation combines icicle plots, visual compression techniques and interactivity, improving space-efficiency and reducing visual structural complexity. We conducted a user study with domain experts to evaluate the usability of OntoPlot, comparing it with the de facto ontology editor Protégé. The results confirm that OntoPlot attains our design goals for association-related tasks and is strongly favoured by domain experts. Michael Wybrow, Yuan-Fang Li, Tobias Czauderna, Yongqun He |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2019 | OSCI: standardized stem cell ontology representation and use cases for stem cell investigationabstractBACKGROUND: Stem cells and stem cell lines are widely used in biomedical research. The Cell Ontology (CL) and Cell Line Ontology (CLO) are two community-based OBO Foundry ontologies in the domains of in vivo cells and in vitro cell line cells, respectively. RESULTS: To support standardized stem cell investigations, we have developed an Ontology for Stem Cell Investigations (OSCI). OSCI imports stem cell and cell line terms from CL and CLO, and investigation-related terms from existing ontologies. A novel focus of OSCI is its application in representing metadata types associated with various stem cell investigations. We also applied OSCI to systematically categorize experimental variables in an induced pluripotent stem cell line cell study related to bipolar disorder. In addition, we used a semi-automated literature mining approach to identify over 200 stem cell gene markers. The relations between these genes and stem cells are modeled and represented in OSCI. CONCLUSIONS: OSCI standardizes stem cells found in vivo and in vitro and in various stem cell investigation processes and entities. The presented use cases demonstrate the utility of OSCI in iPSC studies and literature mining related to bipolar disorder. Yongqun He, William D. Duncan, Daniel J. Cooper, Jens Hansen, Ravi Iyengar, Edison Ong, Kendal Walker, Omar Tibi, Sam Smith, Lucas M. Serra, Jie Zheng 0001, Sirarat Sarntivijai, Stephan C. Schürer, K. Sue O'Shea, Alexander D. Diehl |
BMC Bioinform. | 1 |
| 2019 | A 2018 workshop: vaccine and drug ontology studies (VDOS 2018)abstractThis Editorial first introduces the background of the vaccine and drug relations and how biomedical terminologies and ontologies have been used to support their studies. The history of the seven workshops, initially named VDOSME, and then named VDOS, is also summarized and introduced. Then the 7th International Workshop on Vaccine and Drug Ontology Studies (VDOS 2018), held on August 10th, 2018, Corvallis, Oregon, USA, is introduced in detail. These VDOS workshops have greatly supported the development, applications, and discussion of vaccine- and drug-related terminology and drug studies. Junguk Hur, Cui Tao, Yongqun He |
BMC Bioinform. | 3 |
| 2019 | VIO: ontology classification and study of vaccine responses given various experimental and analytical conditionsabstractBACKGROUND: Different human responses to the same vaccine were frequently observed. For example, independent studies identified overlapping but different transcriptomic gene expression profiles in Yellow Fever vaccine 17D (YF-17D) immunized human subjects. Different experimental and analysis conditions were likely contributed to the observed differences. To investigate this issue, we developed a Vaccine Investigation Ontology (VIO), and applied VIO to classify the different variables and relations among these variables systematically. We then evaluated whether the ontological VIO modeling and VIO-based statistical analysis would contribute to the enhanced vaccine investigation studies and a better understanding of vaccine response mechanisms. RESULTS: Our VIO modeling identified many variables related to data processing and analysis such as normalization method, cut-off criteria, software settings including software version. The datasets from two previous studies on human responses to YF-17D vaccine, reported by Gaucher et al. (2008) and Querec et al. (2009), were re-analyzed. We first applied the same LIMMA statistical method to re-analyze the Gaucher data set and identified a big difference in terms of significantly differentiated gene lists compared to the original study. The different results were likely due to the LIMMA version and software package differences. Our second study re-analyzed both Gaucher and Querec data sets but with the same data processing and analysis pipeline. Significant differences in differential gene lists were also identified. In both studies, we found that Gene Ontology (GO) enrichment results had more overlapping than the gene lists and enriched pathway lists. The visualization of the identified GO hierarchical structures among the enriched GO terms and their associated ancestor terms using GOfox allowed us to find more associations among enriched but often different GO terms, demonstrating the usage of GO hierarchical relations enhance data analysis. CONCLUSIONS: The ontology-based analysis framework supports standardized representation, integration, and analysis of heterogeneous data of host responses to vaccines. Our study also showed that differences in specific variables might explain different results drawn from similar studies. Edison Ong, Peter Sun, Kimberly Berke, Jie Zheng 0001, Guanming Wu, Yongqun He |
BMC Bioinform. | 6 |
| 2019 | The cell line ontology-based representation, integration and analysis of cell lines used in ChinaabstractBACKGROUND: The Chinese National Infrastructure of Cell Line stores and distributes cell lines for biomedical research in China. This study aims to represent and integrate the information of NICR cell lines into the community-based Cell Line Ontology (CLO). RESULTS: We have aligned, represented, and added all identified 2704 cell line cells in NICR to CLO. We also proposed new ontology design patterns to represent the usage of cell line cells as disease models by inducing tumor formation in model organisms, and the relations between cell line cells and their expressed or overexpressed genes or proteins. The resulting CLO-NICR ontology also includes the Chinese representation of the NICR cell line information. CLO-NICR was merged into the general CLO. To serve the cell research community in China, the Chinese version of CLO-NICR was also generated and deposited in the OntoChina ontology repository. The usage of CLO-NICR was demonstrated by DL query and knowledge extraction. CONCLUSIONS: In summary, all identified cell lines from NICR are represented by the semantics framework of CLO and incorporated into CLO as a most recent update. We also generated a CLO-NICR and its Chinese view (CLO-NICR-Cv). The development of CLO-NICR and CLO-NIC-Cv allows the integration of the cell lines from NICR into the community-based CLO ontology and provides an integrative platform to support different applications of CLO in China. Hongjie Pan, Xiaocui Bian, Yongqun He |
BMC Bioinform. | 4 |
| 2019 | Cells in ExperimentaL Life Sciences (CELLS-2018): capturing the knowledge of normal and diseased cells with ontologiesabstractCell cultures and cell lines are widely used in life science experiments. In conjunction with the 2018 International Conference on Biomedical Ontology (ICBO-2018), the 2nd International Workshop on Cells in ExperimentaL Life Science (CELLS-2018) focused on two themes of knowledge representation, for newly-discovered cell types and for cells in disease states. This workshop included five oral presentations and a general discussion session. Two new ontologies, including the Cancer Cell Ontology (CCL) and the Ontology for Stem Cell Investigations (OSCI), were reported in the workshop. In another representation, the Cell Line Ontology (CLO) framework was applied and extended to represent cell line cells used in China and their Chinese representation. Other presentations included a report on the application of ontologies to cross-compare cell types and marker patterns used in flow cytometry studies, and a presentation on new experimental findings about novel cell types based on single cell RNA sequencing assay and their corresponding ontological representation. The general discussion session focused on the ontology design patterns in representing newly-discovered cell types and cells in disease states. Sirarat Sarntivijai, Yongqun He, Alexander D. Diehl |
BMC Bioinform. | 2 |
| 2019 | Machine learning-based identification and rule-based normalization of adverse drug reactions in drug labelsabstractBACKGROUND: Use of medication can cause adverse drug reactions (ADRs), unwanted or unexpected events, which are a major safety concern. Drug labels, or prescribing information or package inserts, describe ADRs. Therefore, systematically identifying ADR information from drug labels is critical in multiple aspects; however, this task is challenging due to the nature of the natural language of drug labels. RESULTS: In this paper, we present a machine learning- and rule-based system for the identification of ADR entity mentions in the text of drug labels and their normalization through the Medical Dictionary for Regulatory Activities (MedDRA) dictionary. The machine learning approach is based on a recently proposed deep learning architecture, which integrates bi-directional Long Short-Term Memory (Bi-LSTM), Convolutional Neural Network (CNN), and Conditional Random Fields (CRF) for entity recognition. The rule-based approach, used for normalizing the identified ADR mentions to MedDRA terms, is based on an extension of our in-house text-mining system, SciMiner. We evaluated our system on the Text Analysis Conference (TAC) Adverse Drug Reaction 2017 challenge test data set, consisting of 200 manually curated US FDA drug labels. Our ML-based system achieved 77.0% F1 score on the task of ADR mention recognition and 82.6% micro-averaged F1 score on the task of ADR normalization, while rule-based system achieved 67.4 and 77.6% F1 scores, respectively. CONCLUSION: Our study demonstrates that a system composed of a deep learning architecture for entity recognition and a rule-based model for entity normalization is a promising approach for ADR extraction from drug labels. Mert Tiftikci, Arzucan Özgür, Yongqun He, Junguk Hur |
BMC Bioinform. | 3 |
| 2019 | ODAE: Ontology-based systematic representation and analysis of drug adverse events and its usage in study of adverse events given different patient age and disease conditionsabstractBACKGROUND: Drug adverse events (AEs), or called adverse drug events (ADEs), are ranked one of the leading causes of mortality. The Ontology of Adverse Events (OAE) has been widely used for adverse event AE representation, standardization, and analysis. OAE-based ADE-specific ontologies, including ODNAE for drug-associated neuropathy-inducing AEs and OCVDAE for cardiovascular drug AEs, have also been developed and used. However, these ADE-specific ontologies do not consider the effects of other factors (e.g., age and drug-treated disease) on the outcomes of ADEs. With more ontological studies of ADEs, it is also critical to develop a general purpose ontology for representing ADEs for various types of drugs. RESULTS: Our survey of FDA drug package insert documents and other resources for 224 neuropathy-inducing drugs discovered that many drugs (e.g., sirolimus and linezolid) cause different AEs given patients' age or the diseases treated by the drugs. To logically represent the complex relations among drug, drug ingredient and mechanism of action, AE, age, disease, and other related factors, an ontology design pattern was developed and applied to generate a community-driven open-source Ontology of Drug Adverse Events (ODAE). The ODAE development follows the OBO Foundry ontology development principles (e.g., openness and collaboration). Built on a generalizable ODAE design pattern and extending the OAE and NDF-RT ontology, ODAE has represented various AEs associated with the over 200 neuropathy-inducing drugs given different age and disease conditions. ODAE is now deposited in the Ontobee for browsing and queries. As a demonstration of usage, a SPARQL query of the ODAE knowledge base was developed to identify all the drugs having the mechanisms of ion channel interactions, the diseases treated with the drugs, and AEs after the treatment in adult patients. AE-specific drug class effects were also explored using ODAE and SPARQL. CONCLUSION: ODAE provides a general representation of ADEs given different conditions and can be used for querying scientific questions. ODAE is also a robust knowledge base and platform for semantic and logic representation and study of ADEs of more drugs in the future. Hong Yu 0014, Solomiya Nysak, Noemi Garg, Edison Ong, Xianwei Ye, Xiangyan Zhang, Yongqun He |
BMC Bioinform. | 7 |
| 2018 | OntoKeeper: Semiotic-driven Ontology Evaluation Tool For Biomedical Ontologists
Muhammad Amith, Frank J. Manion, Chen Liang 0005, Marcelline R. Harris, Dennis Wang, Yongqun He, Cui Tao |
BIBM | 6 |
| 2018 | VRprofile: gene-cluster-detection-based profiling of virulence and antibiotic resistance traits encoded within genome sequences of pathogenic bacteriaabstractVRprofile is a Web server that facilitates rapid investigation of virulence and antibiotic resistance genes, as well as extends these trait transfer-related genetic contexts, in newly sequenced pathogenic bacterial genomes. The used backend database MobilomeDB was firstly built on sets of known gene cluster loci of bacterial type III/IV/VI/VII secretion systems and mobile genetic elements, including integrative and conjugative elements, prophages, class I integrons, IS elements and pathogenicity/antibiotic resistance islands. VRprofile is thus able to co-localize the homologs of these conserved gene clusters using HMMer or BLASTp searches. With the integration of the homologous gene cluster search module with a sequence composition module, VRprofile has exhibited better performance for island-like region predictions than the other widely used methods. In addition, VRprofile also provides an integrated Web interface for aligning and visualizing identified gene clusters with MobilomeDB-archived gene clusters, or a variety set of bacterial genomes. VRprofile might contribute to meet the increasing demands of re-annotations of bacterial variable regions, and aid in the real-time definitions of disease-relevant gene clusters in pathogenic bacteria of interest. VRprofile is freely available at http://bioinfo-mml.sjtu.edu.cn/VRprofile. Cui Tai, Zixin Deng, Weihong Zhong, Yongqun He, Hong-Yu Ou |
Briefings Bioinform. | 5 |
| 2017 | From Scoping Review to Metadata: An Evidence-Based Approach to Identifying Informed Consent Metadata for Biorepositories
Marcelline R. Harris, Frank J. Manion, Hsing-yi Song, Yongqun He, Muhamamd F. Amith, Cui Tao |
AMIA | 4 |
| 2017 | Towards precision informatics of pharmacovigilance: OAE-CTCAE mapping and OAE-based representation and analysis of adverse events in patients treated with cancer drugs
Meiu Wong, Rebecca Racz, Edison Ong, Yongqun He |
AMIA | 4 |
| 2017 | Comparison, alignment, and synchronization of cell line information between CLO and EFOabstractBACKGROUND: The Experimental Factor Ontology (EFO) is an application ontology driven by experimental variables including cell lines to organize and describe the diverse experimental variables and data resided in the EMBL-EBI resources. The Cell Line Ontology (CLO) is an OBO community-based ontology that contains information of immortalized cell lines and relevant experimental components. EFO integrates and extends ontologies from the bio-ontology community to drive a number of practical applications. It is desirable that the community shares design patterns and therefore that EFO reuses the cell line representation from the Cell Line Ontology (CLO). There are, however, challenges to be addressed when developing a common ontology design pattern for representing cell lines in both EFO and CLO. RESULTS: In this study, we developed a strategy to compare and map cell line terms between EFO and CLO. We examined Cellosaurus resources for EFO-CLO cross-references. Text labels of cell lines from both ontologies were verified by biological information axiomatized in each source. The study resulted in the identification 873 EFO-CLO aligned and 344 EFO unique immortalized permanent cell lines. All of these cell lines were updated to CLO and the cell line related information was merged. A design pattern that integrates EFO and CLO was also developed. CONCLUSION: Our study compared, aligned, and synchronized the cell line information between CLO and EFO. The final updated CLO will be examined as the candidate ontology to import and replace eligible EFO cell line classes thereby supporting the interoperability in the bio-ontology domain. Our mapping pipeline illustrates the use of ontology in aiding biological data standardization and integration through the biological and semantics content of cell lines. Edison Ong, Sirarat Sarntivijai, Simon Jupp, Helen E. Parkinson, Yongqun He |
BMC Bioinform. | 5 |
| 2017 | Ontological representation, integration, and analysis of LINCS cell line cells and their cellular responsesabstractBACKGROUND: Aiming to understand cellular responses to different perturbations, the NIH Common Fund Library of Integrated Network-based Cellular Signatures (LINCS) program involves many institutes and laboratories working on over a thousand cell lines. The community-based Cell Line Ontology (CLO) is selected as the default ontology for LINCS cell line representation and integration. RESULTS: CLO has consistently represented all 1097 LINCS cell lines and included information extracted from the LINCS Data Portal and ChEMBL. Using MCF 10A cell line cells as an example, we demonstrated how to ontologically model LINCS cellular signatures such as their non-tumorigenic epithelial cell type, three-dimensional growth, latrunculin-A-induced actin depolymerization and apoptosis, and cell line transfection. A CLO subset view of LINCS cell lines, named LINCS-CLOview, was generated to support systematic LINCS cell line analysis and queries. In summary, LINCS cell lines are currently associated with 43 cell types, 131 tissues and organs, and 121 cancer types. The LINCS-CLO view information can be queried using SPARQL scripts. CONCLUSIONS: CLO was used to support ontological representation, integration, and analysis of over a thousand LINCS cell line cells and their cellular responses. Edison Ong, Jiangan Xie, Zhaohui Ni, Qingping Liu, Sirarat Sarntivijai, Daniel J. Cooper, Raymond Terryn, Vasileios Stathias, Caty Chung, Stephan C. Schürer, Yongqun He |
BMC Bioinform. | 12 |
| 2017 | Cells in experimental life sciences - challenges and solution to the rapid evolution of knowledgeabstractCell cultures used in biomedical experiments come in the form of both sample biopsy primary cells, and maintainable immortalised cell lineages. The rise of bioinformatics and high-throughput technologies has led us to the requirement of ontology representation of cell types and cell lines. The Cell Ontology (CL) and Cell Line Ontology (CLO) have long been established as reference ontologies in the OBO framework. We have compiled a series of the challenges and the proposals of solutions in this CELLS (Cells in ExperimentaL Life Sciences) thematic series that cover the grounds of standing issues and the directions, which were discussed in the First International Workshop on CELLS at the the International Conference on Biomedical Ontology (ICBO). This workshop focused on the extension of the current CL and CLO to cover a wider set of biological questions and challenges needing semantic infrastructure for information modeling. We discussed data-driven use cases that leverage linkage of CL, CLO and other bio-ontologies. This is an established approach in data-driven ontologies such as the Experimental Factor Ontology (EFO), and the Ontology for Biomedical Investigation (OBI). The First International Workshop on CELLS at the International Conference on Biomedical Ontology has brought together experimental biologists and biomedical ontologists to discuss solutions to organizing and representing the rapidly evolving knowledge gained from experimental cells. The workshop has successfully identified the areas of challenge, and the gap in connecting the two domains of knowledge. The outcome of this workshop yielded practical implementation plans to filled in this gap.This CELLS workshop also provided a venue for panel discussions of innovative solutions as well as challenges in the development and applications of biomedical ontologies to represent and analyze experimental cell data. Sirarat Sarntivijai, Alexander D. Diehl, Yongqun He |
BMC Bioinform. | 3 |
| 2016 | Ontology-Based Analysis of Adverse Drug Reactions Associated with Anti-infection Drugs in China
Yuying Cao, Liwei Wang 0010, Yongqun He |
AMIA | 5 |
| 2016 | Application of Ontobedia on Cell Line Ontology Development
Edison Ong, Yongqun He |
AMIA | 2 |
| 2015 | A domain ontology for the Non-Coding RNA fieldabstractIdentification of non-coding RNAs (ncRNAs) has been significantly enhanced due to the rapid advancement in sequencing technologies. On the other hand, semantic annotation of ncRNA data lag behind their identification, and there is a great need to effectively integrate discovery from relevant communities. To this end, the Non-Coding RNA Ontology (NCRO) is being developed to provide a precisely defined ncRNA controlled vocabulary, which can fill a specific and highly needed niche in unification of ncRNA biology. Jingshan Huang, Karen Eilbeck, Judith A. Blake, Dejing Dou, Darren A. Natale, Alan Ruttenberg, Barry Smith 0001, Michael T. Zimmermann, Guoqian Jiang, Bin Wu 0008, Yongqun He, Shaojie Zhang 0001, Xiaowei Wang 0006, Zixing Liu |
BIBM | 12 |
| 2014 | Analysis of Content Coverage for Informed Consent Concepts
Frank J. Manion, Elizabeth Eisenhauer, Alla Karnovsky, Yongqun He, Marcelline R. Harris |
AMIA | 4 |
| 2014 | Ontodog: a web-based ontology community view generation toolabstractBiomedical ontologies are often very large and complex. Only a subset of the ontology may be needed for a specified application or community. For ontology end users, it is desirable to have community-based labels rather than the labels generated by ontology developers. Ontodog is a web-based system that can generate an ontology subset based on Excel input, and support generation of an ontology community view, which is defined as the whole or a subset of the source ontology with user-specified annotations including user-preferred labels. Ontodog allows users to easily generate community views with minimal ontology knowledge and no programming skills or installation required. Currently >100 ontologies including all OBO Foundry ontologies are available to generate the views based on user needs. We demonstrate the application of Ontodog for the generation of community views using the Ontology for Biomedical Investigations as the source ontology. Jie Zheng 0001, Zuoshuang Xiang, Christian J. Stoeckert Jr., Yongqun He |
Bioinform. | 4 |
| 2014 | ICoVax 2013: The 3rd ISV Pre-conference Computational Vaccinology WorkshopabstractFollowing last year's computational vaccinology workshop in Shanghai, China, the third ISV Pre-conference Computational Vaccinology Workshop (ICoVax 2013) was held in Barcelona, Spain. ICoVax 2013 provided an international platform for the attendees to showcase their research and discuss problems and solutions in the development and application of computational vaccinology and vaccine informatics tools. The first of the three full-length papers presented at ICoVax discussed the discovery of viral "camouflage" through cross-conservation of T-cell epitopes using a tool called JanusMatrix. This important paper reports that viruses may camouflage their presence in the human body by incorporating sequences in their proteins that are highly cross-conserved at the T-cell receptor surface with human genome proteins, a discovery that has wide ranging implications for the development of vaccines against viruses that use the camouflage method. The other papers described a database for storing experimentally verified data on DNA vaccines and compared therapeutic targets of western drugs to Chinese herbal medicines for cardiovascular diseases. The short poster presentations covered various uses of informatics tools for processing the DNA and microRNA of pathogens to improve vaccine coverage, efficacy and development. A live (on-line) demonstration of the vaccine design toolkit, iVax, presented by Frances Terry of EpiVax, illustrated how computational vaccinology could be used in the design of next generation vaccines. Anne S. De Groot, Phoebe De Groot, Yongqun He |
BMC Bioinform. | 3 |
| 2014 | DNAVaxDB: the first web-based DNA vaccine database and its data analysisabstractSince the first DNA vaccine studies were done in the 1990s, thousands more studies have followed. Here we report the development and analysis of DNAVaxDB (http://www.violinet.org/dnavaxdb), the first publically available web-based DNA vaccine database that curates, stores, and analyzes experimentally verified DNA vaccines, DNA vaccine plasmid vectors, and protective antigens used in DNA vaccines. All data in DNAVaxDB are annotated from reliable resources, particularly peer-reviewed articles. Among over 140 DNA vaccine plasmids, some plasmids were more frequently used in one type of pathogen than others; for example, pCMVi-UB for G- bacterial DNA vaccines, and pCAGGS for viral DNA vaccines. Presently, over 400 DNA vaccines containing over 370 protective antigens from over 90 infectious and non-infectious diseases have been curated in DNAVaxDB. While extracellular and bacterial cell surface proteins and adhesin proteins were frequently used for DNA vaccine development, the majority of protective antigens used in Chlamydophila DNA vaccines are localized to the inner portion of the cell. The DNA vaccine priming, other vaccine boosting vaccination regimen has been widely used to induce protection against infection of different pathogens such as HIV. Parasitic and cancer DNA vaccines were also systematically analyzed. User-friendly web query and visualization interfaces are available in DNAVaxDB for interactive data search. To support data exchange, the information of DNA vaccines, plasmids, and protective antigens is stored in the Vaccine Ontology (VO). DNAVaxDB is targeted to become a timely and vital source of DNA vaccines and related data and facilitate advanced DNA vaccine research and development. Rebecca Racz, Xinna Li, Mukti Patel, Zuoshuang Xiang, Yongqun He |
BMC Bioinform. | 5 |
| 2013 | Computational vaccinology and the ICoVax 2012 workshopabstractComputational vaccinology or vaccine informatics is an interdisciplinary field that addresses scientific and clinical questions in vaccinology using computational and informatics approaches. Computational vaccinology overlaps with many other fields such as immunoinformatics, reverse vaccinology, postlicensure vaccine research, vaccinomics, literature mining, and systems vaccinology. The second ISV Pre-conference Computational Vaccinology Workshop (ICoVax 2012) was held on October 13, 2013 in Shanghai, China. A number of topics were presented in the workshop, including allergen predictions, prediction of linear T cell epitopes and functional conformational epitopes, prediction of protein-ligand binding regions, vaccine design using reverse vaccinology, and case studies in computational vaccinology. Although a significant progress has been made to date, a number of challenges still exist in the field. This Editorial provides a list of major challenges for the future of computational vaccinology and identifies developing themes that will expand and evolve over the next few years. Yongqun He, Anne S. De Groot, Vladimir Brusic, Christian Schönbach, Nikolai Petrovsky |
BMC Bioinform. | 1 |
| 2013 | Meta-analysis of variables affecting mouse protection efficacy of whole organism Brucella vaccines and vaccine candidatesabstractBACKGROUND: Vaccine protection investigation includes three processes: vaccination, pathogen challenge, and vaccine protection efficacy assessment. Many variables can affect the results of vaccine protection. Brucella, a genus of facultative intracellular bacteria, is the etiologic agent of brucellosis in humans and multiple animal species. Extensive research has been conducted in developing effective live attenuated Brucella vaccines. We hypothesized that some variables play a more important role than others in determining vaccine protective efficacy. Using Brucella vaccines and vaccine candidates as study models, this hypothesis was tested by meta-analysis of Brucella vaccine studies reported in the literature. RESULTS: Nineteen variables related to vaccine-induced protection of mice against infection with virulent brucellae were selected based on modeling investigation of the vaccine protection processes. The variable "vaccine protection efficacy" was set as a dependent variable while the other eighteen were set as independent variables. Discrete or continuous values were collected from papers for each variable of each data set. In total, 401 experimental groups were manually annotated from 74 peer-reviewed publications containing mouse protection data for live attenuated Brucella vaccines or vaccine candidates. Our ANOVA analysis indicated that nine variables contributed significantly (P-value < 0.05) to Brucella vaccine protection efficacy: vaccine strain, vaccination host (mouse) strain, vaccination dose, vaccination route, challenge pathogen strain, challenge route, challenge-killing interval, colony forming units (CFUs) in mouse spleen, and CFU reduction compared to control group. The other 10 variables (e.g., mouse age, vaccination-challenge interval, and challenge dose) were not found to be statistically significant (P-value > 0.05). The protection level of RB51 was sacrificed when the values of several variables (e.g., vaccination route, vaccine viability, and challenge pathogen strain) change. It is suggestive that it is difficult to protect against aerosol challenge. Somewhat counter-intuitively, our results indicate that intraperitoneal and subcutaneous vaccinations are much more effective to protect against aerosol Brucella challenge than intranasal vaccination. CONCLUSIONS: Literature meta-analysis identified variables that significantly contribute to Brucella vaccine protection efficacy. The results obtained provide critical information for rational vaccine study design. Literature meta-analysis is generic and can be applied to analyze variables critical for vaccine protection against other infectious diseases. Thomas E. Todd, Omar Tibi, Samantha Sayers, Denise N. Bronner, Zuoshuang Xiang, Yongqun He |
BMC Bioinform. | 7 |
| 2013 | Genome-wide prediction of vaccine targets for human herpes simplex viruses using Vaxign reverse vaccinologyabstractHerpes simplex virus (HSV) types 1 and 2 (HSV-1 and HSV-2) are the most common infectious agents of humans. No safe and effective HSV vaccines have been licensed. Reverse vaccinology is an emerging and revolutionary vaccine development strategy that starts with the prediction of vaccine targets by informatics analysis of genome sequences. Vaxign (http://www.violinet.org/vaxign) is the first web-based vaccine design program based on reverse vaccinology. In this study, we used Vaxign to analyze 52 herpesvirus genomes, including 3 HSV-1 genomes, one HSV-2 genome, 8 other human herpesvirus genomes, and 40 non-human herpesvirus genomes. The HSV-1 strain 17 genome that contains 77 proteins was used as the seed genome. These 77 proteins are conserved in two other HSV-1 strains (strain F and strain H129). Two envelope glycoproteins gJ and gG do not have orthologs in HSV-2 or 8 other human herpesviruses. Seven HSV-1 proteins (including gJ and gG) do not have orthologs in all 40 non-human herpesviruses. Nineteen proteins are conserved in all human herpesviruses, including capsid scaffold protein UL26.5 (NP_044628.1). As the only HSV-1 protein predicted to be an adhesin, UL26.5 is a promising vaccine target. The MHC Class I and II epitopes were predicted by the Vaxign Vaxitop prediction program and IEDB prediction programs recently installed and incorporated in Vaxign. Our comparative analysis found that the two programs identified largely the same top epitopes but also some positive results predicted from one program might not be positive from another program. Overall, our Vaxign computational prediction provides many promising candidates for rational HSV vaccine development. The method is generic and can also be used to predict other viral vaccine targets. Zuoshuang Xiang, Yongqun He |
BMC Bioinform. | 2 |
| 2007 | miniTUBA: medical inference by network integration of temporal data using Bayesian analysisabstractMOTIVATION: Many biomedical and clinical research problems involve discovering causal relationships between observations gathered from temporal events. Dynamic Bayesian networks are a powerful modeling approach to describe causal or apparently causal relationships, and support complex medical inference, such as future response prediction, automated learning, and rational decision making. Although many engines exist for creating Bayesian networks, most require a local installation and significant data manipulation to be practical for a general biologist or clinician. No software pipeline currently exists for interpretation and inference of dynamic Bayesian networks learned from biomedical and clinical data. RESULTS: miniTUBA is a web-based modeling system that allows clinical and biomedical researchers to perform complex medical/clinical inference and prediction using dynamic Bayesian network analysis with temporal datasets. The software allows users to choose different analysis parameters (e.g. Markov lags and prior topology), and continuously update their data and refine their results. miniTUBA can make temporal predictions to suggest interventions based on an automated learning process pipeline using all data provided. Preliminary tests using synthetic data and laboratory research data indicate that miniTUBA accurately identifies regulatory network structures from temporal data. AVAILABILITY: miniTUBA is available at http://www.minituba.org. Zuoshuang Xiang, Rebecca M. Minter, Xiaoming Bi, Peter J. Woolf, Yongqun He |
Bioinform. | 5 |
| 2007 | CRCView: a web server for analyzing and visualizing microarray gene expression data using model-based clusteringabstractUNLABELLED: CRCView is a user-friendly point-and-click web server for analyzing and visualizing microarray gene expression data using a Dirichlet process mixture model-based clustering algorithm. CRCView is designed to clustering genes based on their expression profiles. It allows flexible input data format, rich graphical illustration as well as integrated GO term based annotation/interpretation of clustering results. AVAILABILITY: http://helab.bioinformatics.med.umich.edu/crcview/. Zuoshuang Xiang, Zhaohui S. Qin, Yongqun He |
Bioinform. | 3 |
| 2006 | BBP: Brucella genome annotation with literature mining and curationabstractBACKGROUND: Brucella species are Gram-negative, facultative intracellular bacteria that cause brucellosis in humans and animals. Sequences of four Brucella genomes have been published, and various Brucella gene and genome data and analysis resources exist. A web gateway to integrate these resources will greatly facilitate Brucella research. Brucella genome data in current databases is largely derived from computational analysis without experimental validation typically found in peer-reviewed publications. It is partially due to the lack of a literature mining and curation system able to efficiently incorporate the large amount of literature data into genome annotation. It is further hypothesized that literature-based Brucella gene annotation would increase understanding of complicated Brucella pathogenesis mechanisms. RESULTS: The Brucella Bioinformatics Portal (BBP) is developed to integrate existing Brucella genome data and analysis tools with literature mining and curation. The BBP InterBru database and Brucella Genome Browser allow users to search and analyze genes of 4 currently available Brucella genomes and link to more than 20 existing databases and analysis programs. Brucella literature publications in PubMed are extracted and can be searched by a TextPresso-powered natural language processing method, a MeSH browser, a keywords search, and an automatic literature update service. To efficiently annotate Brucella genes using the large amount of literature publications, a literature mining and curation system coined Limix is developed to integrate computational literature mining methods with a PubSearch-powered manual curation and management system. The Limix system is used to quickly find and confirm 107 Brucella gene mutations including 75 genes shown to be essential for Brucella virulence. The 75 genes are further clustered using COG. In addition, 62 Brucella genetic interactions are extracted from literature publications. These results make possible more comprehensive investigation of Brucella pathogenesis. Other BBP features include publication email alert service, Brucella researchers' contact database, and discussion forum. CONCLUSION: BBP is a gateway for Brucella researchers to search, analyze, and curate Brucella genome data originated from public databases and literature. Brucella gene mutations and genetic interactions are annotated using Limix leading to better understanding of Brucella pathogenesis. Zuoshuang Xiang, Yongqun He |
BMC Bioinform. | 3 |
| 2005 | PIML: the Pathogen Information Markup LanguageabstractMOTIVATION: A vast amount of information about human, animal and plant pathogens has been acquired, stored and displayed in varied formats through different resources, both electronically and otherwise. However, there is no community standard format for organizing this information or agreement on machine-readable format(s) for data exchange, thereby hampering interoperation efforts across information systems harboring such infectious disease data. RESULTS: The Pathogen Information Markup Language (PIML) is a free, open, XML-based format for representing pathogen information. XSLT-based visual presentations of valid PIML documents were developed and can be accessed through the PathInfo website or as part of the interoperable web services federation known as ToolBus/PathPort. Currently, detailed PIML documents are available for 21 pathogens deemed of high priority with regard to public health and national biological defense. A dynamic query system allows simple queries as well as comparisons among these pathogens. Continuing efforts are being taken to include other groups' supporting PIML and to develop more PIML documents. AVAILABILITY: All the PIML-related information is accessible from http://www.vbi.vt.edu/pathport/pathinfo/ Yongqun He, Richard R. Vines, Alice R. Wattam, Georgiy V. Abramochkin, Allan Dickerman, J. Dana Eckart, Bruno W. S. Sobral |
Bioinform. | 1 |