Qian Zhu 0003

dblp:02/562-3 · DBLP profile ↗
← Back
39ranked-venue papers
14as first author
16since 2021 · last 2024
0000-0002-4858-6333ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 13 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
YearPublicationVenuePosition
2024 Aligning Orphanet Classification to Identify Disease Characteristics among Rare Disease Clusters
abstract
Understanding the underlying etiologies of rare diseases may facilitate research across multiple conditions, enabling basket trail design and drug repurposing. In this study, we aligned clusters of rare diseases with Orphanet classifications to represent their shared etiologies and establish a foundation for further investigation on underly biological mechanism discovery. By utilizing the linearized Orphanet categories, we connected 35 clusters of rare diseases into 18 classifications. Significant associations were found between the categories "Rare Developmental Defects During Embryogenesis" and "Rare Inborn Errors of Metabolism" and the clusters in this study, suggesting that many rare diseases originating in the prenatal period or related to metabolism may present a substantial opportunity for success in future investigation.
Sungrim Moon, Jessica Maine, Ewy A. Mathé, Qian Zhu 0003
BIBM4
2024 Identifying Drug Repurposing Candidates for CLN3 Targeting Proteomics Expression Profile
abstract
Juvenile neuronal ceroid lipofuscinosis (CLN3) is a rare neurodegenerative disorder lacking effective therapies. This study aimed at developing a drug repurposing approach to identify potential therapeutic candidates for CLN3 using its protein expression profile (CPEP) constructed from proteomics data. Differentially expressed proteins were identified and applied to query the iLINCS database, resulting in 60 FDA-approved drugs with reversal effects on CPEP. These candidates were further prioritized based on regulation strength, coverage, and blood-brain barrier permeability. Top candidates include Vorinostat and Cyclosporine, which have shown promise due to their significant regulation scores and blood-brain barrier permeation probability. These results provide opportunities for further investigation on novel therapies for CLN3.
Shixue Sun, Rosemary Mejia, An N. Dang Do, Qian Zhu 0003
BIBM4
2024 An application of studying FAERS data to Enhance Drug Safety and Treatment Outcomes in Rare Diseases
abstract
Rare diseases affect fewer than 200,000 individuals in the United States, with some being so rare that only a handful of people are impacted. According to the U.S. Food and Drug Administration (FDA), there are 1,268 approved orphan drugs available for treating these conditions. However, potentially beneficial drugs can also have side effects. Some adverse events, while serious, may be rare, making them difficult to identify or quantify in randomized controlled trials. Understanding these events is critical for improving patient safety and treatment outcomes. To better assess these risks, we aimed at summarizing adverse drug events for rare diseases by utilizing FDA Adverse Event Reporting System (FAERS). This study offers a foundation for future research of improving drug safety in rare diseases.
Jaber Valinejad, Yanji Xu, Qian Zhu 0003
BIBM3
2023 Mining NIH BTRIS Data for Drug Repurposing: A Case Study of Glioblastoma
abstract
The purpose of drug repurposing is to identify alternative uses of FDA approved drugs, which significantly accelerates the drug development process. Meanwhile, clinical data illustrate the patterns and clinical outcomes of drug use, so they have been increasingly applied to support drug development, particularly for drug repurposing. The NIH Biomedical Translational Research Information System (BTRIS) is a resource which compiles deidentified patient data from clinical research done across NIH Institutes and Centers. In this study, we analyzed clinical data available from BTRIS to identify drug repurposing candidates, i.e., identifying drugs that were correlated with an increased survival rate for glioblastoma (GBM) patients. Specifically, we extracted all the administered drugs on GBM patients and fitted them to elastic-net penalized Cox proportional hazards (CPH) models, a regression model for investigating the association between the survival rate of patients and covariates (administered drugs in this study). We were able to identify several potential drug candidates for GBM to be further evaluated with other data types and by performing biological experiments.
Shixue Sun, Yitao Tian, Qian Zhu 0003
BIBM3
2023 Clinical Assessment of Pneumocystosis with MIMIC Data
abstract
Pneumocystosis remains a life-threatening disease with a high mortality rate. It's critical to understand its clinical course and risk factors for better disease management. In this retrospective analysis, we aimed to elucidate the prognostic determinants of in-hospital mortality among patients diagnosed with pneumocystosis. Data were extracted from the Medical Information Mart for Intensive Care (MIMIC)-IV database, encompassing all recorded cases of pneumocystosis. The dataset included patient admission records, comprehensive laboratory results, and medication administration data, which were meticulously analyzed to identify relevant features. Employing logistic regression and random forest, we discerned that the administration of micafungin sodium and vasopressin have significant impacts as risk factors on the survival rate of pneumocystosis patients.
Huanfei Wang, Qian Zhu 0003, Jian Pei 0001
BIBM2
2023 Prediction of Drug Targets based on In Vitro Activity Profiles Toward Drug Repurposing for Rare Diseases
abstract
Over 300 million people are suffering from rare diseases, most of which have limited treatment options. Therefore, discovering new treatments for rare diseases is imperative. Drug repurposing, which identifies new uses for approved drugs, is considered one of the viable and risk-managed strategies for disease treatments. To promote the drug repurposing process, we introduced a prediction model to uncover novel relationships between gene targets and chemical compounds. In our previous study, we identified enriched genes for compounds from the Toxicology in the 21st Century program (Tox21) 10K library, to extend that study for enriched gene target prediction, we developed machine learning (ML) models including Support Vector Machine; K-Nearest Neighbors; Random Forest; and extreme gradient boosting (XGBoost), by using Tox21 bioassay screening data. All four models perform well with f1_score over 0.7, and XGBoost has the best performance with four different multi-label prediction embedding algorithms, including Binary Relevance; Label Powerset; Classifier Chain; Multi-Output Classifier. Our study explored a reliable method to predict potential gene targets from in vitro activity profile data toward drug repurposing.
Binghan Xue, Ruili Huang, Qian Zhu 0003, Yanji Xu
BIBM3
2023 Dual projection learning with adaptive graph smoothing for multi-label classification
Rui-hang Cai, Timothy Apasiba Abeo, Qian Zhu 0003, Cong-hua Zhou, Xiangjun Shen
Appl. Intell.4
2023 Multi-dictionary induced low-rank representation with multi-manifold regularization
Jinghui Zhou, Xiangjun Shen, Sixing Liu, Liangjun Wang, Qian Zhu 0003, Ping Qian
Appl. Intell.5
2023 Clustering rare diseases within an ontology-enriched knowledge graph
abstract
OBJECTIVE: Identifying sets of rare diseases with shared aspects of etiology and pathophysiology may enable drug repurposing. Toward that aim, we utilized an integrative knowledge graph to construct clusters of rare diseases. MATERIALS AND METHODS: Data on 3242 rare diseases were extracted from the National Center for Advancing Translational Science Genetic and Rare Diseases Information center internal data resources. The rare disease data enriched with additional biomedical data, including gene and phenotype ontologies, biological pathway data, and small molecule-target activity data, to create a knowledge graph (KG). Node embeddings were trained and clustered. We validated the disease clusters through semantic similarity and feature enrichment analysis. RESULTS: Thirty-seven disease clusters were created with a mean size of 87 diseases. We validate the clusters quantitatively via semantic similarity based on the Orphanet Rare Disease Ontology. In addition, the clusters were analyzed for enrichment of associated genes, revealing that the enriched genes within clusters are highly related. DISCUSSION: We demonstrate that node embeddings are an effective method for clustering diseases within a heterogenous KG. Semantically similar diseases and relevant enriched genes have been uncovered within the clusters. Connections between disease clusters and drugs are enumerated for follow-up efforts. CONCLUSION: We lay out a method for clustering rare diseases using graph node embeddings. We develop an easy-to-maintain pipeline that can be updated when new data on rare diseases emerges. The embeddings themselves can be paired with other representation learning methods for other data types, such as drugs, to address other predictive modeling problems.
Jaleal Sanjak, Jessica Binder, Arjun Singh Yadaw, Qian Zhu 0003, Ewy A. Mathé
J. Am. Medical Informatics Assoc.4
2023 Multi-view Representation Induced Kernel Ensemble Support Vector Machine
Ebenezer Quayson, Ernest Domanaanmwi Ganaa, Qian Zhu 0003, Xiangjun Shen
Neural Process. Lett.3
2022 Profiling Tox21 Bioassay Towards Drug Repurposing for Rare Diseases
Yanji Xu, Jaleal Sanjak, Andrew Patt, Ruili Huang, Chunxu Qu, Qian Zhu 0003
AMIA7
2022 Semantic Annotation of NIH Funding Data for Supporting Rare Disease Research
abstract
With the advances in science and technology, the number of research in rare diseases has dramatically increased over the past twenty years. Systematically accessing those research projects funded by NIH would allow us to assess the current status of research, and research gaps remain in this area. Consequently, new research might be inspired to bridge the gaps. We previously developed a knowledge graph to semantically represent NIH funded rare disease research projects by analyzing project titles. To expand the use of NIH funding data, in this study we extended the previous work in two folds, 1) we applied our self-developed NLP package named NormMap to identify rare disease related projects, 2) we semantically annotated project titles and abstracts with biomedical concepts in UMLS to illustrate the project aims. With such rich information extracted from NIH funding data via semantic annotation, an updated version of the knowledge graph will be developed to advance rare disease research as the next step.
Szeling Hsu, Sue Qu, Yanji Xu, Qian Zhu 0003
BIBM4
2022 Integrative Rare Disease Profile Creation via NormMap to Advance Rare Disease Research
abstract
Given the nature of rare diseases, lack of data and standards impedes research in rare diseases. A method to improve data interoperability is necessary to allow data reuse, integration, and exchange in rare disease. A computational package named NormMap was developed to identify rare disease related data from various types of resources in free text via semantic annotation with rare disease terms from NCATS Genetic and Rare Diseases (GARD). In this preliminary study, four different sources which include NIH funded projects, clinical trials, PubMed articles, and Reddit subreddits, were applied to generate rare disease profiles by extending and exploring NormMap. Those profiles would offer a complete view of rare diseases from different aspects, funding agencies, patient groups, scientific research, to ultimately advance rare disease research, which is demonstrated in our case study.
Devon Leadman, Yanji Xu, Sue Qu, Qian Zhu 0003
BIBM4
2021 Data Normalization Improves Semantic Annotation - a Case Study of Rare Disease Name Annotation
abstract
Despite the individually low prevalence of rare diseases, they collectively constitute a big challenge to human health. Accurate annotation of rare diseases in biomedical data through natural language processing (NLP) could be instrumental in biomedical informatics research. Herein, we propose a data normalization-based annotation approach, complementary to the popular biomedical data annotation tool named MetaMap, for disease annotation from biomedical data in free text.
Charlie Tang, Yanji Xu, Qian Zhu 0003
BIBM3
2021 Scientific Evidence Based Knowledge Graph in Rare Diseases
abstract
Rare diseases are naturally associated with low prevalence rate, which raises a big challenge due to less data available for supporting preclinical and clinical studies. Therefore, it is critical to fully utilize the accumulated scientific publications in rare diseases over years, in order to access full spectrum of scientific research and enable relevant scientific evidence extraction and generation. In this study, we obtained rare disease related PubMed articles, extracted multiple types of biomedical information, and semantically presented the data in a knowledge graph, which is hosted in Neo4j based on a predefined data model to support further rare disease research.
Qian Zhu 0003, Ruizheng Liu, Gunjan Vatas, Andrew Clough, Yanji Xu, Dac-Trung Nguyen, Ewy A. Mathé, Eric Sid
BIBM1
2021 Better Understand Rare Disease Patients' Needs by Analyzing Social Media Data - a Case Study of Cystic Fibrosis
abstract
There are approximately 7,000 rare diseases, and 25-30 million people affected with a rare disease in the United States. Prevalence rate for rare diseases is relatively low compared to common diseases. Thus, disease rarity leads to a lack of clinical familiarity that often impedes accurate and timely diagnosis for many rare disease patients. Social media has become an important resource/tool for discussing, sharing, and seeking information relevant to rare diseases by patients and families. In this study, we aimed to analyze rare disease-related posts from Reddit, one popular social media platform to reveal rare disease patients' needs based on hidden topics to be identified. We implemented NLP/topic modelling to identify main topics from the posts and consequently computed TF-IDF to detect the most prevalent phrases as sub-topics from the posts. As a proof of concept, we primarily focused on Cystic Fibrosis as a case study to demonstrate the use of data from Reddit for rare disease research.
Qian Zhu 0003, Eric Sundstrom, Yanji Xu
BIBM1
2019 Sex, obesity, diabetes, and exposure to particulate matter among patients with severe asthma: Scientific insights from a comparative analysis of open clinical data sources during a five-day hackathon
abstract
This special communication describes activities, products, and lessons learned from a recent hackathon that was funded by the National Center for Advancing Translational Sciences via the Biomedical Data Translator program ('Translator'). Specifically, Translator team members self-organized and worked together to conceptualize and execute, over a five-day period, a multi-institutional clinical research study that aimed to examine, using open clinical data sources, relationships between sex, obesity, diabetes, and exposure to airborne fine particulate matter among patients with severe asthma. The goal was to develop a proof of concept that this new model of collaboration and data sharing could effectively produce meaningful scientific results and generate new scientific hypotheses. Three Translator Clinical Knowledge Sources, each of which provides open access (via Application Programming Interfaces) to data derived from the electronic health record systems of major academic institutions, served as the source of study data. Jupyter Python notebooks, shared in GitHub repositories, were used to call the knowledge sources and analyze and integrate the results. The results replicated established or suspected relationships between sex, obesity, diabetes, exposure to airborne fine particulate matter, and severe asthma. In addition, the results demonstrated specific differences across the three Translator Clinical Knowledge Sources, suggesting cohort- and/or environment-specific factors related to the services themselves or the catchment area from which each service derives patient data. Collectively, this special communication demonstrates the power and utility of intense, team-oriented hackathons and offers general technical, organizational, and scientific lessons learned.
Karamarie Fecho, Stanley C. Ahalt, Saravanan Arunachalam, James Champion, Christopher G. Chute, Sarah Davis, Kenneth Gersing, Gwênlyn Glusman, Jennifer Hadlock, Jewel Lee, Emily R. Pfaff, Max Robinson, Eric Sid, Casey N. Ta, Hao Xu 0006, Richard L. Zhu, Qian Zhu 0003, David B. Peden
J. Biomed. Informatics17
2019 Multi-layer framework of identifying placenta related research towards Placenta Curated Research Dataset (PCRD) development for the PAT project
Qian Zhu 0003, Shanna Frierson, Alicia Francis, Lydia Rogers, Daniel Lyman
J. Biomed. Informatics1
2016 Analyzing and retrieving illicit drug-related posts from social media
abstract
Illicit drug use is a serious problem around the world. Social media has increasingly become an important tool for analyzing drug use patterns and monitoring emerging drug abuse trends. Accurately retrieving illicit drug-related social media posts is an important step in this research. Frequently, hashtags are used to identify and retrieve posts on a specific topic. However hashtags are highly ambiguous. Posts with the same hashtags are not always on the same topic. Moreover, hashtags are evolving, especially those related to illicit drugs. New street names are introduced constantly to avoid detection. In this paper, we employ topic modeling to disambiguate hashtags and track the changes of hashtags using semantic word embedding. Our preliminary evaluation shows the promise of these methods.
Tao Ding 0007, Arpita Roy, Zhiyuan Chen 0003, Qian Zhu 0003, Shimei Pan
BIBM4
2016 Risk feature assessment of readmission for diabetes
abstract
About 382 million people have Diabetes in 2013, and the International Diabetes Federation estimated that there are 4.9 million people died from Diabetes in 2014. Diabetes continues to be a chronic disease plagued by frequent hospital readmissions. In order to better understand the risk features impacting readmissions for future prevention and management, in this study, we programmatically analyzed a large clinical dataset containing more than 100,000 clinical records for diabetes patients from 130 US hospitals. Specifically, we developed three different machine learning algorithms, Logistic Regression, Random Forest and manipulated Random Forest to identify and prioritize the most significant risk features. By comparing the results generated by these three methods, the manipulated Random Forest illustrates greater capacity of generating a more complete and concrete list of readmission related risk features. Such method is generalizable and can be applied in other disease oriented studies.
Qian Zhu 0003, Anirudh Akkati, Pornpoh Hongwattanakul
BIBM1
2015 Acquisition of diabetes-related biological associations using a motif based network: Preliminary results
abstract
Diabetes remains the 7th leading cause of death in the US in 2010. There is an urgent need to find effective solutions to prevent and treat diabetes. With the advance in computational technology, computer-aided approaches has attracted much interest given the promising findings been generated accordingly. In this study, we introduced an approach that applied network analysis to reveal novel biological associations for diabetes from a large biological interaction network. In particular, network motif discovery has been performed to identify diabetes relevant motif-based network, from where the diabetes relevant biological associations can be detected via perturbation. The associations identified from this study will not only illustrate possible biological mechanism for diabetes, but also support biomedical application development ultimately, such as drug repositioning.
Iyanuoluwa Emmanuel Odebode, Aryya Gangopadhyay, Qian Zhu 0003
BIBM3
2015 Desiderata for computable representations of electronic health records-driven phenotype algorithms
abstract
BACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages.
Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris
J. Am. Medical Informatics Assoc.7
2015 Review and evaluation of electronic health records-driven phenotype algorithm authoring tools for clinical and translational research
abstract
OBJECTIVE: To review and evaluate available software tools for electronic health record-driven phenotype authoring in order to identify gaps and needs for future development. MATERIALS AND METHODS: Candidate phenotype authoring tools were identified through (1) literature search in four publication databases (PubMed, Embase, Web of Science, and Scopus) and (2) a web search. A collection of tools was compiled and reviewed after the searches. A survey was designed and distributed to the developers of the reviewed tools to discover their functionalities and features. RESULTS: Twenty-four different phenotype authoring tools were identified and reviewed. Developers of 16 of these identified tools completed the evaluation survey (67% response rate). The surveyed tools showed commonalities but also varied in their capabilities in algorithm representation, logic functions, data support and software extensibility, search functions, user interface, and data outputs. DISCUSSION: Positive trends identified in the evaluation included: algorithms can be represented in both computable and human readable formats; and most tools offer a web interface for easy access. However, issues were also identified: many tools were lacking advanced logic functions for authoring complex algorithms; the ability to construct queries that leveraged un-structured data was not widely implemented; and many tools had limited support for plug-ins or external analytic software. CONCLUSIONS: Existing phenotype authoring tools could enable clinical researchers to work with electronic health record data more efficiently, but gaps still exist in terms of the functionalities of such tools. The present work can serve as a reference point for the future development of similar tools.
Jie Xu 0011, Luke V. Rasmussen, Pamela L. Shaw, Guoqian Jiang, Richard C. Kiefer, Huan Mo, Jennifer A. Pacheco, Peter Speltz, Qian Zhu 0003, Joshua C. Denny, Jyotishman Pathak, William K. Thompson, Enid N. H. Montague
J. Am. Medical Informatics Assoc.9
2015 Toward a complete dataset of drug-drug interaction information from publicly available sources
abstract
Although potential drug-drug interactions (PDDIs) are a significant source of preventable drug-related harm, there is currently no single complete source of PDDI information. In the current study, all publically available sources of PDDI information that could be identified using a comprehensive and broad search were combined into a single dataset. The combined dataset merged fourteen different sources including 5 clinically-oriented information sources, 4 Natural Language Processing (NLP) Corpora, and 5 Bioinformatics/Pharmacovigilance information sources. As a comprehensive PDDI source, the merged dataset might benefit the pharmacovigilance text mining community by making it possible to compare the representativeness of NLP corpora for PDDI text extraction tasks, and specifying elements that can be useful for future PDDI extraction purposes. An analysis of the overlap between and across the data sources showed that there was little overlap. Even comprehensive PDDI lists such as DrugBank, KEGG, and the NDF-RT had less than 50% overlap with each other. Moreover, all of the comprehensive lists had incomplete coverage of two data sources that focus on PDDIs of interest in most clinical settings. Based on this information, we think that systems that provide access to the comprehensive lists, such as APIs into RxNorm, should be careful to inform users that the lists may be incomplete with respect to PDDIs that drug experts suggest clinicians be aware of. In spite of the low degree of overlap, several dozen cases were identified where PDDI information provided in drug product labeling might be augmented by the merged dataset. Moreover, the combined dataset was also shown to improve the performance of an existing PDDI NLP pipeline and a recently published PDDI pharmacovigilance protocol. Future work will focus on improvement of the methods for mapping between PDDI information sources, identifying methods to improve the use of the merged dataset in PDDI NLP algorithms, integrating high-quality PDDI information from the merged dataset into Wikidata, and making the combined dataset accessible as Semantic Web Linked Data.
Serkan Ayvaz, John R. Horn, Oktie Hassanzadeh, Qian Zhu 0003, Johann Stan, Nicholas P. Tatonetti, Santiago Vilar, Mathias Brochhausen, Matthias Samwald, Majid Rastegar-Mojarad, Michel Dumontier, Richard D. Boyce
J. Biomed. Informatics4
2014 Evaluation of RxNorm for Medication Clinical Decision Support
Robert R. Freimuth, Kelly Wix, Qian Zhu 0003, Mark Siska, Christopher G. Chute
AMIA3
2014 Evaluation of Existing Phenotype Authoring Tools for Clinical Research
Luke V. Rasmussen, Jie Xu 0011, Ruijue Liu, Qian Zhu 0003, Jennifer A. Pacheco, Jyotishman Pathak, William K. Thompson, Joshua C. Denny, Huan Mo, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague
AMIA4
2014 iGenetics: An Individualized Genetic Test Recommendation System Based on EHRs
Qian Zhu 0003, Christopher G. Chute, Matthew Ferber
AMIA1
2014 Qualitative evaluation of three phenotype information models to find methotrexate liver injury
Qian Zhu 0003, Huan Mo, Luke V. Rasmussen, Andrew R. Post, Jennifer A. Pacheco, Jie Xu 0011, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Jyotishman Pathak
AMIA1
2014 Genetic testing knowledge base (GTKB) towards individualized genetic test recommendation - An experimental study
abstract
The gap between a large growing number of genetic tests and a suboptimal clinical workflow of incorporating these tests into regular clinical practice poses barriers to effective reliance on advanced genetic technologies to improve quality of healthcare. A promising solution to encourage and assist physicians to incorporate genetic tests in their clinical practice is an intelligent genetic test recommendation system for 1) providing a comprehensive view of genetic tests as education resources; 2) recommending the most appropriate genetic tests to patients based on clinical evidence. In this paper, we introduce a genetic testing knowledge base, called GTKB, which was designed to support further individualized genetic test recommendation. More specifically, we extracted clinical characteristics identified from Electronic Health Records (EHRs) that have been used as phenotypic information for linked archived biological material to accelerate research in individualized medicine, and well-documented public genetic testing resources including Genetic Testing Registry (GTR) and published genetic testing guidelines (GTG) to construct a genetic test orientated knowledge base, ultimately supporting genetic test recommendation. An experimental study for “wilson disease mutation screen test” has been conducted to demonstrate the identification of salient clinical characteristics and the process of incorporating EHR derived phenotypes into the GTKB construction.
Qian Zhu 0003, Christopher G. Chute, Matthew Ferber
BIBM1
2014 Pharmacological class data representation in the Web Ontology Language (OWL)
abstract
Dozens of drug terminologies and resources capture drug and/or drug class information; they range greatly in their coverage and their adequacy of representation. However, there are no transformative ways to link these resources together in a standard and formal fashion, which hinders data integration and data representation for supporting drug-related clinical and translational studies. In this study, we generated a standardized Pharmacological Class Profile Ontology, named PCPO, which integrates multiple drug resources in the Web Ontology Language (OWL). More specifically, we mapped two well-known drug class resources, Anatomical Therapeutic Chemical classification system (ATC) and National Drug File Reference Terminology (NDF-RT), as the pharmacological class backbone. Furthermore we extended the PCPO with individual clinical drug information extracted from RxNorm and Structured Product Labeling. In parallel, we calculated and compared chemical structure similarity for each drug pair from ATC and NDF-RT, and re-grouped drugs with a similar structure into a same drug class. PCPO will not only present big drug data into well-organized formal form, OWL with possible inference capability, but also potentially support computational drug repurposing application development.
Qian Zhu 0003, Cui Tao
IEEE BigData1
2013 Using standardized clinical data modeling and knowledge representation to compute pharmacogenomic data elements
Qian Zhu 0003, Jyotishman Pathak, Robert R. Freimuth, Christopher G. Chute
AMIA1
2013 A semantic-web oriented representation of the clinical element model for secondary use of electronic health records data
abstract
The clinical element model (CEM) is an information model designed for representing clinical information in electronic health records (EHR) systems across organizations. The current representation of CEMs does not support formal semantic definitions and therefore it is not possible to perform reasoning and consistency checking on derived models. This paper introduces our efforts to represent the CEM specification using the Web Ontology Language (OWL). The CEM-OWL representation connects the CEM content with the Semantic Web environment, which provides authoring, reasoning, and querying tools. This work may also facilitate the harmonization of the CEMs with domain knowledge represented in terminology models as well as other clinical information models such as the openEHR archetype model. We have created the CEM-OWL meta ontology based on the CEM specification. A convertor has been implemented in Java to automatically translate detailed CEMs from XML to OWL. A panel evaluation has been conducted, and the results show that the OWL modeling can faithfully represent the CEM specification and represent patient data.
Cui Tao, Guoqian Jiang, Thomas A. Oniki, Robert R. Freimuth, Qian Zhu 0003, Deepak K. Sharma, Jyotishman Pathak, Stanley M. Huff, Christopher G. Chute
J. Am. Medical Informatics Assoc.5
2013 Harmonization and semantic annotation of data dictionaries from the Pharmacogenomics Research Network: A case study
Qian Zhu 0003, Robert R. Freimuth, Zonghui Lian, Scott Bauer, Jyotishman Pathak, Cui Tao, Matthew J. Durski, Christopher G. Chute
J. Biomed. Informatics1
2013 Disambiguation of PharmGKB drug-disease relations with NDF-RT and SPL
Qian Zhu 0003, Robert R. Freimuth, Jyotishman Pathak, Matthew J. Durski, Christopher G. Chute
J. Biomed. Informatics1
2012 A Standardized Drug and Drug Class Universal Network
Qian Zhu 0003, Guoqian Jiang, Christopher G. Chute
AMIA1
2011 Semantic Inference using Chemogenomics Data for Drug Discovery
abstract
BACKGROUND: Semantic Web Technology (SWT) makes it possible to integrate and search the large volume of life science datasets in the public domain, as demonstrated by well-known linked data projects such as LODD, Bio2RDF, and Chem2Bio2RDF. Integration of these sets creates large networks of information. We have previously described a tool called WENDI for aggregating information pertaining to new chemical compounds, effectively creating evidence paths relating the compounds to genes, diseases and so on. In this paper we examine the utility of automatically inferring new compound-disease associations (and thus new links in the network) based on semantically marked-up versions of these evidence paths, rule-sets and inference engines. RESULTS: Through the implementation of a semantic inference algorithm, rule set, Semantic Web methods (RDF, OWL and SPARQL) and new interfaces, we have created a new tool called Chemogenomic Explorer that uses networks of ontologically annotated RDF statements along with deductive reasoning tools to infer new associations between the query structure and genes and diseases from WENDI results. The tool then permits interactive clustering and filtering of these evidence paths. CONCLUSIONS: We present a new aggregate approach to inferring links between chemical compounds and diseases using semantic inference. This approach allows multiple evidence paths between compounds and diseases to be identified using a rule-set and semantically annotated data, and for these evidence paths to be clustered to show overall evidence linking the compound to a disease. We believe this is a powerful approach, because it allows compound-disease relationships to be ranked by the amount of evidence supporting them.
Qian Zhu 0003, Yuyin Sun, Sashikiran Challa, Ying Ding 0001, Michael S. Lajiness, David J. Wild 0001
BMC Bioinform.1
2010 Chem2Bio2RDF: A Linked Open Data Portal for Systems Chemical Biology
abstract
The Chem2Bio2RDF portal is a Linked Open Data (LOD) portal for systems chemical biology aiming for facilitating drug discovery. It converts around 25 different datasets on genes, compounds, drugs, pathways, side effects, diseases, and MEDLINE/PubMed documents into RDF triples and links them to other LOD bubbles, such as Bio2RDF, LODD and DBPedia. The portal is based on D2R server and provides a SPARQL endpoint, but adds on several unique features such as RDF faceted browser, user-friendly SPARQL query generator, MEDLINE/PubMed cross validation service, and Cytoscape visualization plugin. Three use cases demonstrate the functionality and usability of this portal. The portal is available at http://chem2bio2rdf.org.
Bin Chen 0002, Ying Ding 0001, David J. Wild 0001, Yuyin Sun, Qian Zhu 0003, Madhuvanthi Sankaranarayanan
Web Intelligence7
2010 Using Web Technologies for Integrative Drug Discovery
abstract
Recent years have seen a huge increase in the amount of publicly-available information relevant to drug discovery, including online databases of compound and bioassay information; scholarly publications linking compounds with genes, targets and diseases; and predictive models that can suggest new links between compounds, genes, targets and diseases. However, there is a lack of tools and methods to integrate this information, and in particular to look for pertinent knowledge and relationships across multiple sources. At Indiana University we are tackling this problem by applying aggregative data mining tools and semantic web technologies including using an extensive web service infrastructure, RDF networks and inference engines, ontologies, and automated extraction of information from scholarly literature.
Qian Zhu 0003, Sashikiran Challa, Prajakta Purohit, Yuyin Sun, Michael S. Lajiness, David J. Wild 0001, Ying Ding 0001
Web Intelligence1
2010 Chem2Bio2RDF: a semantic framework for linking and data mining chemogenomic and systems chemical biology data
abstract
BACKGROUND: Recently there has been an explosion of new data sources about genes, proteins, genetic variations, chemical compounds, diseases and drugs. Integration of these data sources and the identification of patterns that go across them is of critical interest. Initiatives such as Bio2RDF and LODD have tackled the problem of linking biological data and drug data respectively using RDF. Thus far, the inclusion of chemogenomic and systems chemical biology information that crosses the domains of chemistry and biology has been very limited RESULTS: We have created a single repository called Chem2Bio2RDF by aggregating data from multiple chemogenomics repositories that is cross-linked into Bio2RDF and LODD. We have also created a linked-path generation tool to facilitate SPARQL query generation, and have created extended SPARQL functions to address specific chemical/biological search needs. We demonstrate the utility of Chem2Bio2RDF in investigating polypharmacology, identification of potential multiple pathway inhibitors, and the association of pathways with adverse drug reactions. CONCLUSIONS: We have created a new semantic systems chemical biology resource, and have demonstrated its potential usefulness in specific examples of polypharmacology, multiple pathway inhibition and adverse drug reaction--pathway mapping. We have also demonstrated the usefulness of extending SPARQL with cheminformatics and bioinformatics functionality.
Bin Chen 0002, Dazhi Jiao, Qian Zhu 0003, Ying Ding 0001, David J. Wild 0001
BMC Bioinform.5