EDBT 2026 Demo / reviewers in the wild / expert
Qian Zhu 0003
dblp:02/562-3
· DBLP profile ↗
39ranked-venue papers
14as first author
16since 2021 · last 2024
0000-0002-4858-6333ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 13 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Aligning Orphanet Classification to Identify Disease Characteristics among Rare Disease ClustersabstractUnderstanding the underlying etiologies of rare diseases may facilitate research across multiple conditions, enabling basket trail design and drug repurposing. In this study, we aligned clusters of rare diseases with Orphanet classifications to represent their shared etiologies and establish a foundation for further investigation on underly biological mechanism discovery. By utilizing the linearized Orphanet categories, we connected 35 clusters of rare diseases into 18 classifications. Significant associations were found between the categories "Rare Developmental Defects During Embryogenesis" and "Rare Inborn Errors of Metabolism" and the clusters in this study, suggesting that many rare diseases originating in the prenatal period or related to metabolism may present a substantial opportunity for success in future investigation. Sungrim Moon, Jessica Maine, Ewy A. Mathé, Qian Zhu 0003 |
BIBM | 4 |
| 2024 | Identifying Drug Repurposing Candidates for CLN3 Targeting Proteomics Expression ProfileabstractJuvenile neuronal ceroid lipofuscinosis (CLN3) is a rare neurodegenerative disorder lacking effective therapies. This study aimed at developing a drug repurposing approach to identify potential therapeutic candidates for CLN3 using its protein expression profile (CPEP) constructed from proteomics data. Differentially expressed proteins were identified and applied to query the iLINCS database, resulting in 60 FDA-approved drugs with reversal effects on CPEP. These candidates were further prioritized based on regulation strength, coverage, and blood-brain barrier permeability. Top candidates include Vorinostat and Cyclosporine, which have shown promise due to their significant regulation scores and blood-brain barrier permeation probability. These results provide opportunities for further investigation on novel therapies for CLN3. Shixue Sun, Rosemary Mejia, An N. Dang Do, Qian Zhu 0003 |
BIBM | 4 |
| 2024 | An application of studying FAERS data to Enhance Drug Safety and Treatment Outcomes in Rare DiseasesabstractRare diseases affect fewer than 200,000 individuals in the United States, with some being so rare that only a handful of people are impacted. According to the U.S. Food and Drug Administration (FDA), there are 1,268 approved orphan drugs available for treating these conditions. However, potentially beneficial drugs can also have side effects. Some adverse events, while serious, may be rare, making them difficult to identify or quantify in randomized controlled trials. Understanding these events is critical for improving patient safety and treatment outcomes. To better assess these risks, we aimed at summarizing adverse drug events for rare diseases by utilizing FDA Adverse Event Reporting System (FAERS). This study offers a foundation for future research of improving drug safety in rare diseases. Jaber Valinejad, Yanji Xu, Qian Zhu 0003 |
BIBM | 3 |
| 2023 | Mining NIH BTRIS Data for Drug Repurposing: A Case Study of GlioblastomaabstractThe purpose of drug repurposing is to identify alternative uses of FDA approved drugs, which significantly accelerates the drug development process. Meanwhile, clinical data illustrate the patterns and clinical outcomes of drug use, so they have been increasingly applied to support drug development, particularly for drug repurposing. The NIH Biomedical Translational Research Information System (BTRIS) is a resource which compiles deidentified patient data from clinical research done across NIH Institutes and Centers. In this study, we analyzed clinical data available from BTRIS to identify drug repurposing candidates, i.e., identifying drugs that were correlated with an increased survival rate for glioblastoma (GBM) patients. Specifically, we extracted all the administered drugs on GBM patients and fitted them to elastic-net penalized Cox proportional hazards (CPH) models, a regression model for investigating the association between the survival rate of patients and covariates (administered drugs in this study). We were able to identify several potential drug candidates for GBM to be further evaluated with other data types and by performing biological experiments. Shixue Sun, Yitao Tian, Qian Zhu 0003 |
BIBM | 3 |
| 2023 | Clinical Assessment of Pneumocystosis with MIMIC DataabstractPneumocystosis remains a life-threatening disease with a high mortality rate. It's critical to understand its clinical course and risk factors for better disease management. In this retrospective analysis, we aimed to elucidate the prognostic determinants of in-hospital mortality among patients diagnosed with pneumocystosis. Data were extracted from the Medical Information Mart for Intensive Care (MIMIC)-IV database, encompassing all recorded cases of pneumocystosis. The dataset included patient admission records, comprehensive laboratory results, and medication administration data, which were meticulously analyzed to identify relevant features. Employing logistic regression and random forest, we discerned that the administration of micafungin sodium and vasopressin have significant impacts as risk factors on the survival rate of pneumocystosis patients. Huanfei Wang, Qian Zhu 0003, Jian Pei 0001 |
BIBM | 2 |
| 2023 | Prediction of Drug Targets based on In Vitro Activity Profiles Toward Drug Repurposing for Rare DiseasesabstractOver 300 million people are suffering from rare diseases, most of which have limited treatment options. Therefore, discovering new treatments for rare diseases is imperative. Drug repurposing, which identifies new uses for approved drugs, is considered one of the viable and risk-managed strategies for disease treatments. To promote the drug repurposing process, we introduced a prediction model to uncover novel relationships between gene targets and chemical compounds. In our previous study, we identified enriched genes for compounds from the Toxicology in the 21st Century program (Tox21) 10K library, to extend that study for enriched gene target prediction, we developed machine learning (ML) models including Support Vector Machine; K-Nearest Neighbors; Random Forest; and extreme gradient boosting (XGBoost), by using Tox21 bioassay screening data. All four models perform well with f1_score over 0.7, and XGBoost has the best performance with four different multi-label prediction embedding algorithms, including Binary Relevance; Label Powerset; Classifier Chain; Multi-Output Classifier. Our study explored a reliable method to predict potential gene targets from in vitro activity profile data toward drug repurposing. Binghan Xue, Ruili Huang, Qian Zhu 0003, Yanji Xu |
BIBM | 3 |
| 2023 | Dual projection learning with adaptive graph smoothing for multi-label classification
Rui-hang Cai, Timothy Apasiba Abeo, Qian Zhu 0003, Cong-hua Zhou, Xiangjun Shen |
Appl. Intell. | 4 |
| 2023 | Multi-dictionary induced low-rank representation with multi-manifold regularization
Jinghui Zhou, Xiangjun Shen, Sixing Liu, Liangjun Wang, Qian Zhu 0003, Ping Qian |
Appl. Intell. | 5 |
| 2023 | Clustering rare diseases within an ontology-enriched knowledge graphabstractOBJECTIVE: Identifying sets of rare diseases with shared aspects of etiology and pathophysiology may enable drug repurposing. Toward that aim, we utilized an integrative knowledge graph to construct clusters of rare diseases. MATERIALS AND METHODS: Data on 3242 rare diseases were extracted from the National Center for Advancing Translational Science Genetic and Rare Diseases Information center internal data resources. The rare disease data enriched with additional biomedical data, including gene and phenotype ontologies, biological pathway data, and small molecule-target activity data, to create a knowledge graph (KG). Node embeddings were trained and clustered. We validated the disease clusters through semantic similarity and feature enrichment analysis. RESULTS: Thirty-seven disease clusters were created with a mean size of 87 diseases. We validate the clusters quantitatively via semantic similarity based on the Orphanet Rare Disease Ontology. In addition, the clusters were analyzed for enrichment of associated genes, revealing that the enriched genes within clusters are highly related. DISCUSSION: We demonstrate that node embeddings are an effective method for clustering diseases within a heterogenous KG. Semantically similar diseases and relevant enriched genes have been uncovered within the clusters. Connections between disease clusters and drugs are enumerated for follow-up efforts. CONCLUSION: We lay out a method for clustering rare diseases using graph node embeddings. We develop an easy-to-maintain pipeline that can be updated when new data on rare diseases emerges. The embeddings themselves can be paired with other representation learning methods for other data types, such as drugs, to address other predictive modeling problems. Jaleal Sanjak, Jessica Binder, Arjun Singh Yadaw, Qian Zhu 0003, Ewy A. Mathé |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Multi-view Representation Induced Kernel Ensemble Support Vector Machine
Ebenezer Quayson, Ernest Domanaanmwi Ganaa, Qian Zhu 0003, Xiangjun Shen |
Neural Process. Lett. | 3 |
| 2022 | Profiling Tox21 Bioassay Towards Drug Repurposing for Rare Diseases
Yanji Xu, Jaleal Sanjak, Andrew Patt, Ruili Huang, Chunxu Qu, Qian Zhu 0003 |
AMIA | 7 |
| 2022 | Semantic Annotation of NIH Funding Data for Supporting Rare Disease ResearchabstractWith the advances in science and technology, the number of research in rare diseases has dramatically increased over the past twenty years. Systematically accessing those research projects funded by NIH would allow us to assess the current status of research, and research gaps remain in this area. Consequently, new research might be inspired to bridge the gaps. We previously developed a knowledge graph to semantically represent NIH funded rare disease research projects by analyzing project titles. To expand the use of NIH funding data, in this study we extended the previous work in two folds, 1) we applied our self-developed NLP package named NormMap to identify rare disease related projects, 2) we semantically annotated project titles and abstracts with biomedical concepts in UMLS to illustrate the project aims. With such rich information extracted from NIH funding data via semantic annotation, an updated version of the knowledge graph will be developed to advance rare disease research as the next step. Szeling Hsu, Sue Qu, Yanji Xu, Qian Zhu 0003 |
BIBM | 4 |
| 2022 | Integrative Rare Disease Profile Creation via NormMap to Advance Rare Disease ResearchabstractGiven the nature of rare diseases, lack of data and standards impedes research in rare diseases. A method to improve data interoperability is necessary to allow data reuse, integration, and exchange in rare disease. A computational package named NormMap was developed to identify rare disease related data from various types of resources in free text via semantic annotation with rare disease terms from NCATS Genetic and Rare Diseases (GARD). In this preliminary study, four different sources which include NIH funded projects, clinical trials, PubMed articles, and Reddit subreddits, were applied to generate rare disease profiles by extending and exploring NormMap. Those profiles would offer a complete view of rare diseases from different aspects, funding agencies, patient groups, scientific research, to ultimately advance rare disease research, which is demonstrated in our case study. Devon Leadman, Yanji Xu, Sue Qu, Qian Zhu 0003 |
BIBM | 4 |
| 2021 | Data Normalization Improves Semantic Annotation - a Case Study of Rare Disease Name AnnotationabstractDespite the individually low prevalence of rare diseases, they collectively constitute a big challenge to human health. Accurate annotation of rare diseases in biomedical data through natural language processing (NLP) could be instrumental in biomedical informatics research. Herein, we propose a data normalization-based annotation approach, complementary to the popular biomedical data annotation tool named MetaMap, for disease annotation from biomedical data in free text. Charlie Tang, Yanji Xu, Qian Zhu 0003 |
BIBM | 3 |
| 2021 | Scientific Evidence Based Knowledge Graph in Rare DiseasesabstractRare diseases are naturally associated with low prevalence rate, which raises a big challenge due to less data available for supporting preclinical and clinical studies. Therefore, it is critical to fully utilize the accumulated scientific publications in rare diseases over years, in order to access full spectrum of scientific research and enable relevant scientific evidence extraction and generation. In this study, we obtained rare disease related PubMed articles, extracted multiple types of biomedical information, and semantically presented the data in a knowledge graph, which is hosted in Neo4j based on a predefined data model to support further rare disease research. Qian Zhu 0003, Ruizheng Liu, Gunjan Vatas, Andrew Clough, Yanji Xu, Dac-Trung Nguyen, Ewy A. Mathé, Eric Sid |
BIBM | 1 |
| 2021 | Better Understand Rare Disease Patients' Needs by Analyzing Social Media Data - a Case Study of Cystic FibrosisabstractThere are approximately 7,000 rare diseases, and 25-30 million people affected with a rare disease in the United States. Prevalence rate for rare diseases is relatively low compared to common diseases. Thus, disease rarity leads to a lack of clinical familiarity that often impedes accurate and timely diagnosis for many rare disease patients. Social media has become an important resource/tool for discussing, sharing, and seeking information relevant to rare diseases by patients and families. In this study, we aimed to analyze rare disease-related posts from Reddit, one popular social media platform to reveal rare disease patients' needs based on hidden topics to be identified. We implemented NLP/topic modelling to identify main topics from the posts and consequently computed TF-IDF to detect the most prevalent phrases as sub-topics from the posts. As a proof of concept, we primarily focused on Cystic Fibrosis as a case study to demonstrate the use of data from Reddit for rare disease research. Qian Zhu 0003, Eric Sundstrom, Yanji Xu |
BIBM | 1 |
| 2019 | Sex, obesity, diabetes, and exposure to particulate matter among patients with severe asthma: Scientific insights from a comparative analysis of open clinical data sources during a five-day hackathonabstractThis special communication describes activities, products, and lessons learned from a recent hackathon that was funded by the National Center for Advancing Translational Sciences via the Biomedical Data Translator program ('Translator'). Specifically, Translator team members self-organized and worked together to conceptualize and execute, over a five-day period, a multi-institutional clinical research study that aimed to examine, using open clinical data sources, relationships between sex, obesity, diabetes, and exposure to airborne fine particulate matter among patients with severe asthma. The goal was to develop a proof of concept that this new model of collaboration and data sharing could effectively produce meaningful scientific results and generate new scientific hypotheses. Three Translator Clinical Knowledge Sources, each of which provides open access (via Application Programming Interfaces) to data derived from the electronic health record systems of major academic institutions, served as the source of study data. Jupyter Python notebooks, shared in GitHub repositories, were used to call the knowledge sources and analyze and integrate the results. The results replicated established or suspected relationships between sex, obesity, diabetes, exposure to airborne fine particulate matter, and severe asthma. In addition, the results demonstrated specific differences across the three Translator Clinical Knowledge Sources, suggesting cohort- and/or environment-specific factors related to the services themselves or the catchment area from which each service derives patient data. Collectively, this special communication demonstrates the power and utility of intense, team-oriented hackathons and offers general technical, organizational, and scientific lessons learned. Karamarie Fecho, Stanley C. Ahalt, Saravanan Arunachalam, James Champion, Christopher G. Chute, Sarah Davis, Kenneth Gersing, Gwênlyn Glusman, Jennifer Hadlock, Jewel Lee, Emily R. Pfaff, Max Robinson, Eric Sid, Casey N. Ta, Hao Xu 0006, Richard L. Zhu, Qian Zhu 0003, David B. Peden |
J. Biomed. Informatics | 17 |
| 2019 | Multi-layer framework of identifying placenta related research towards Placenta Curated Research Dataset (PCRD) development for the PAT project
Qian Zhu 0003, Shanna Frierson, Alicia Francis, Lydia Rogers, Daniel Lyman |
J. Biomed. Informatics | 1 |
| 2016 | Analyzing and retrieving illicit drug-related posts from social mediaabstractIllicit drug use is a serious problem around the world. Social media has increasingly become an important tool for analyzing drug use patterns and monitoring emerging drug abuse trends. Accurately retrieving illicit drug-related social media posts is an important step in this research. Frequently, hashtags are used to identify and retrieve posts on a specific topic. However hashtags are highly ambiguous. Posts with the same hashtags are not always on the same topic. Moreover, hashtags are evolving, especially those related to illicit drugs. New street names are introduced constantly to avoid detection. In this paper, we employ topic modeling to disambiguate hashtags and track the changes of hashtags using semantic word embedding. Our preliminary evaluation shows the promise of these methods. Tao Ding 0007, Arpita Roy, Zhiyuan Chen 0003, Qian Zhu 0003, Shimei Pan |
BIBM | 4 |
| 2016 | Risk feature assessment of readmission for diabetesabstractAbout 382 million people have Diabetes in 2013, and the International Diabetes Federation estimated that there are 4.9 million people died from Diabetes in 2014. Diabetes continues to be a chronic disease plagued by frequent hospital readmissions. In order to better understand the risk features impacting readmissions for future prevention and management, in this study, we programmatically analyzed a large clinical dataset containing more than 100,000 clinical records for diabetes patients from 130 US hospitals. Specifically, we developed three different machine learning algorithms, Logistic Regression, Random Forest and manipulated Random Forest to identify and prioritize the most significant risk features. By comparing the results generated by these three methods, the manipulated Random Forest illustrates greater capacity of generating a more complete and concrete list of readmission related risk features. Such method is generalizable and can be applied in other disease oriented studies. Qian Zhu 0003, Anirudh Akkati, Pornpoh Hongwattanakul |
BIBM | 1 |
| 2015 | Acquisition of diabetes-related biological associations using a motif based network: Preliminary resultsabstractDiabetes remains the 7th leading cause of death in the US in 2010. There is an urgent need to find effective solutions to prevent and treat diabetes. With the advance in computational technology, computer-aided approaches has attracted much interest given the promising findings been generated accordingly. In this study, we introduced an approach that applied network analysis to reveal novel biological associations for diabetes from a large biological interaction network. In particular, network motif discovery has been performed to identify diabetes relevant motif-based network, from where the diabetes relevant biological associations can be detected via perturbation. The associations identified from this study will not only illustrate possible biological mechanism for diabetes, but also support biomedical application development ultimately, such as drug repositioning. Iyanuoluwa Emmanuel Odebode, Aryya Gangopadhyay, Qian Zhu 0003 |
BIBM | 3 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 7 |
| 2015 | Review and evaluation of electronic health records-driven phenotype algorithm authoring tools for clinical and translational researchabstractOBJECTIVE: To review and evaluate available software tools for electronic health record-driven phenotype authoring in order to identify gaps and needs for future development. MATERIALS AND METHODS: Candidate phenotype authoring tools were identified through (1) literature search in four publication databases (PubMed, Embase, Web of Science, and Scopus) and (2) a web search. A collection of tools was compiled and reviewed after the searches. A survey was designed and distributed to the developers of the reviewed tools to discover their functionalities and features. RESULTS: Twenty-four different phenotype authoring tools were identified and reviewed. Developers of 16 of these identified tools completed the evaluation survey (67% response rate). The surveyed tools showed commonalities but also varied in their capabilities in algorithm representation, logic functions, data support and software extensibility, search functions, user interface, and data outputs. DISCUSSION: Positive trends identified in the evaluation included: algorithms can be represented in both computable and human readable formats; and most tools offer a web interface for easy access. However, issues were also identified: many tools were lacking advanced logic functions for authoring complex algorithms; the ability to construct queries that leveraged un-structured data was not widely implemented; and many tools had limited support for plug-ins or external analytic software. CONCLUSIONS: Existing phenotype authoring tools could enable clinical researchers to work with electronic health record data more efficiently, but gaps still exist in terms of the functionalities of such tools. The present work can serve as a reference point for the future development of similar tools. Jie Xu 0011, Luke V. Rasmussen, Pamela L. Shaw, Guoqian Jiang, Richard C. Kiefer, Huan Mo, Jennifer A. Pacheco, Peter Speltz, Qian Zhu 0003, Joshua C. Denny, Jyotishman Pathak, William K. Thompson, Enid N. H. Montague |
J. Am. Medical Informatics Assoc. | 9 |
| 2015 | Toward a complete dataset of drug-drug interaction information from publicly available sourcesabstractAlthough potential drug-drug interactions (PDDIs) are a significant source of preventable drug-related harm, there is currently no single complete source of PDDI information. In the current study, all publically available sources of PDDI information that could be identified using a comprehensive and broad search were combined into a single dataset. The combined dataset merged fourteen different sources including 5 clinically-oriented information sources, 4 Natural Language Processing (NLP) Corpora, and 5 Bioinformatics/Pharmacovigilance information sources. As a comprehensive PDDI source, the merged dataset might benefit the pharmacovigilance text mining community by making it possible to compare the representativeness of NLP corpora for PDDI text extraction tasks, and specifying elements that can be useful for future PDDI extraction purposes. An analysis of the overlap between and across the data sources showed that there was little overlap. Even comprehensive PDDI lists such as DrugBank, KEGG, and the NDF-RT had less than 50% overlap with each other. Moreover, all of the comprehensive lists had incomplete coverage of two data sources that focus on PDDIs of interest in most clinical settings. Based on this information, we think that systems that provide access to the comprehensive lists, such as APIs into RxNorm, should be careful to inform users that the lists may be incomplete with respect to PDDIs that drug experts suggest clinicians be aware of. In spite of the low degree of overlap, several dozen cases were identified where PDDI information provided in drug product labeling might be augmented by the merged dataset. Moreover, the combined dataset was also shown to improve the performance of an existing PDDI NLP pipeline and a recently published PDDI pharmacovigilance protocol. Future work will focus on improvement of the methods for mapping between PDDI information sources, identifying methods to improve the use of the merged dataset in PDDI NLP algorithms, integrating high-quality PDDI information from the merged dataset into Wikidata, and making the combined dataset accessible as Semantic Web Linked Data. Serkan Ayvaz, John R. Horn, Oktie Hassanzadeh, Qian Zhu 0003, Johann Stan, Nicholas P. Tatonetti, Santiago Vilar, Mathias Brochhausen, Matthias Samwald, Majid Rastegar-Mojarad, Michel Dumontier, Richard D. Boyce |
J. Biomed. Informatics | 4 |
| 2014 | Evaluation of RxNorm for Medication Clinical Decision Support
Robert R. Freimuth, Kelly Wix, Qian Zhu 0003, Mark Siska, Christopher G. Chute |
AMIA | 3 |
| 2014 | Evaluation of Existing Phenotype Authoring Tools for Clinical Research
Luke V. Rasmussen, Jie Xu 0011, Ruijue Liu, Qian Zhu 0003, Jennifer A. Pacheco, Jyotishman Pathak, William K. Thompson, Joshua C. Denny, Huan Mo, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague |
AMIA | 4 |
| 2014 | iGenetics: An Individualized Genetic Test Recommendation System Based on EHRs
Qian Zhu 0003, Christopher G. Chute, Matthew Ferber |
AMIA | 1 |
| 2014 | Qualitative evaluation of three phenotype information models to find methotrexate liver injury
Qian Zhu 0003, Huan Mo, Luke V. Rasmussen, Andrew R. Post, Jennifer A. Pacheco, Jie Xu 0011, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Jyotishman Pathak |
AMIA | 1 |
| 2014 | Genetic testing knowledge base (GTKB) towards individualized genetic test recommendation - An experimental studyabstractThe gap between a large growing number of genetic tests and a suboptimal clinical workflow of incorporating these tests into regular clinical practice poses barriers to effective reliance on advanced genetic technologies to improve quality of healthcare. A promising solution to encourage and assist physicians to incorporate genetic tests in their clinical practice is an intelligent genetic test recommendation system for 1) providing a comprehensive view of genetic tests as education resources; 2) recommending the most appropriate genetic tests to patients based on clinical evidence. In this paper, we introduce a genetic testing knowledge base, called GTKB, which was designed to support further individualized genetic test recommendation. More specifically, we extracted clinical characteristics identified from Electronic Health Records (EHRs) that have been used as phenotypic information for linked archived biological material to accelerate research in individualized medicine, and well-documented public genetic testing resources including Genetic Testing Registry (GTR) and published genetic testing guidelines (GTG) to construct a genetic test orientated knowledge base, ultimately supporting genetic test recommendation. An experimental study for “wilson disease mutation screen test” has been conducted to demonstrate the identification of salient clinical characteristics and the process of incorporating EHR derived phenotypes into the GTKB construction. Qian Zhu 0003, Christopher G. Chute, Matthew Ferber |
BIBM | 1 |
| 2014 | Pharmacological class data representation in the Web Ontology Language (OWL)abstractDozens of drug terminologies and resources capture drug and/or drug class information; they range greatly in their coverage and their adequacy of representation. However, there are no transformative ways to link these resources together in a standard and formal fashion, which hinders data integration and data representation for supporting drug-related clinical and translational studies. In this study, we generated a standardized Pharmacological Class Profile Ontology, named PCPO, which integrates multiple drug resources in the Web Ontology Language (OWL). More specifically, we mapped two well-known drug class resources, Anatomical Therapeutic Chemical classification system (ATC) and National Drug File Reference Terminology (NDF-RT), as the pharmacological class backbone. Furthermore we extended the PCPO with individual clinical drug information extracted from RxNorm and Structured Product Labeling. In parallel, we calculated and compared chemical structure similarity for each drug pair from ATC and NDF-RT, and re-grouped drugs with a similar structure into a same drug class. PCPO will not only present big drug data into well-organized formal form, OWL with possible inference capability, but also potentially support computational drug repurposing application development. Qian Zhu 0003, Cui Tao |
IEEE BigData | 1 |
| 2013 | Using standardized clinical data modeling and knowledge representation to compute pharmacogenomic data elements
Qian Zhu 0003, Jyotishman Pathak, Robert R. Freimuth, Christopher G. Chute |
AMIA | 1 |
| 2013 | A semantic-web oriented representation of the clinical element model for secondary use of electronic health records dataabstractThe clinical element model (CEM) is an information model designed for representing clinical information in electronic health records (EHR) systems across organizations. The current representation of CEMs does not support formal semantic definitions and therefore it is not possible to perform reasoning and consistency checking on derived models. This paper introduces our efforts to represent the CEM specification using the Web Ontology Language (OWL). The CEM-OWL representation connects the CEM content with the Semantic Web environment, which provides authoring, reasoning, and querying tools. This work may also facilitate the harmonization of the CEMs with domain knowledge represented in terminology models as well as other clinical information models such as the openEHR archetype model. We have created the CEM-OWL meta ontology based on the CEM specification. A convertor has been implemented in Java to automatically translate detailed CEMs from XML to OWL. A panel evaluation has been conducted, and the results show that the OWL modeling can faithfully represent the CEM specification and represent patient data. Cui Tao, Guoqian Jiang, Thomas A. Oniki, Robert R. Freimuth, Qian Zhu 0003, Deepak K. Sharma, Jyotishman Pathak, Stanley M. Huff, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | Harmonization and semantic annotation of data dictionaries from the Pharmacogenomics Research Network: A case study
Qian Zhu 0003, Robert R. Freimuth, Zonghui Lian, Scott Bauer, Jyotishman Pathak, Cui Tao, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2013 | Disambiguation of PharmGKB drug-disease relations with NDF-RT and SPL
Qian Zhu 0003, Robert R. Freimuth, Jyotishman Pathak, Matthew J. Durski, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2012 | A Standardized Drug and Drug Class Universal Network
Qian Zhu 0003, Guoqian Jiang, Christopher G. Chute |
AMIA | 1 |
| 2011 | Semantic Inference using Chemogenomics Data for Drug DiscoveryabstractBACKGROUND: Semantic Web Technology (SWT) makes it possible to integrate and search the large volume of life science datasets in the public domain, as demonstrated by well-known linked data projects such as LODD, Bio2RDF, and Chem2Bio2RDF. Integration of these sets creates large networks of information. We have previously described a tool called WENDI for aggregating information pertaining to new chemical compounds, effectively creating evidence paths relating the compounds to genes, diseases and so on. In this paper we examine the utility of automatically inferring new compound-disease associations (and thus new links in the network) based on semantically marked-up versions of these evidence paths, rule-sets and inference engines. RESULTS: Through the implementation of a semantic inference algorithm, rule set, Semantic Web methods (RDF, OWL and SPARQL) and new interfaces, we have created a new tool called Chemogenomic Explorer that uses networks of ontologically annotated RDF statements along with deductive reasoning tools to infer new associations between the query structure and genes and diseases from WENDI results. The tool then permits interactive clustering and filtering of these evidence paths. CONCLUSIONS: We present a new aggregate approach to inferring links between chemical compounds and diseases using semantic inference. This approach allows multiple evidence paths between compounds and diseases to be identified using a rule-set and semantically annotated data, and for these evidence paths to be clustered to show overall evidence linking the compound to a disease. We believe this is a powerful approach, because it allows compound-disease relationships to be ranked by the amount of evidence supporting them. Qian Zhu 0003, Yuyin Sun, Sashikiran Challa, Ying Ding 0001, Michael S. Lajiness, David J. Wild 0001 |
BMC Bioinform. | 1 |
| 2010 | Chem2Bio2RDF: A Linked Open Data Portal for Systems Chemical BiologyabstractThe Chem2Bio2RDF portal is a Linked Open Data (LOD) portal for systems chemical biology aiming for facilitating drug discovery. It converts around 25 different datasets on genes, compounds, drugs, pathways, side effects, diseases, and MEDLINE/PubMed documents into RDF triples and links them to other LOD bubbles, such as Bio2RDF, LODD and DBPedia. The portal is based on D2R server and provides a SPARQL endpoint, but adds on several unique features such as RDF faceted browser, user-friendly SPARQL query generator, MEDLINE/PubMed cross validation service, and Cytoscape visualization plugin. Three use cases demonstrate the functionality and usability of this portal. The portal is available at http://chem2bio2rdf.org. Bin Chen 0002, Ying Ding 0001, David J. Wild 0001, Yuyin Sun, Qian Zhu 0003, Madhuvanthi Sankaranarayanan |
Web Intelligence | 7 |
| 2010 | Using Web Technologies for Integrative Drug DiscoveryabstractRecent years have seen a huge increase in the amount of publicly-available information relevant to drug discovery, including online databases of compound and bioassay information; scholarly publications linking compounds with genes, targets and diseases; and predictive models that can suggest new links between compounds, genes, targets and diseases. However, there is a lack of tools and methods to integrate this information, and in particular to look for pertinent knowledge and relationships across multiple sources. At Indiana University we are tackling this problem by applying aggregative data mining tools and semantic web technologies including using an extensive web service infrastructure, RDF networks and inference engines, ontologies, and automated extraction of information from scholarly literature. Qian Zhu 0003, Sashikiran Challa, Prajakta Purohit, Yuyin Sun, Michael S. Lajiness, David J. Wild 0001, Ying Ding 0001 |
Web Intelligence | 1 |
| 2010 | Chem2Bio2RDF: a semantic framework for linking and data mining chemogenomic and systems chemical biology dataabstractBACKGROUND: Recently there has been an explosion of new data sources about genes, proteins, genetic variations, chemical compounds, diseases and drugs. Integration of these data sources and the identification of patterns that go across them is of critical interest. Initiatives such as Bio2RDF and LODD have tackled the problem of linking biological data and drug data respectively using RDF. Thus far, the inclusion of chemogenomic and systems chemical biology information that crosses the domains of chemistry and biology has been very limited RESULTS: We have created a single repository called Chem2Bio2RDF by aggregating data from multiple chemogenomics repositories that is cross-linked into Bio2RDF and LODD. We have also created a linked-path generation tool to facilitate SPARQL query generation, and have created extended SPARQL functions to address specific chemical/biological search needs. We demonstrate the utility of Chem2Bio2RDF in investigating polypharmacology, identification of potential multiple pathway inhibitors, and the association of pathways with adverse drug reactions. CONCLUSIONS: We have created a new semantic systems chemical biology resource, and have demonstrated its potential usefulness in specific examples of polypharmacology, multiple pathway inhibition and adverse drug reaction--pathway mapping. We have also demonstrated the usefulness of extending SPARQL with cheminformatics and bioinformatics functionality. Bin Chen 0002, Dazhi Jiao, Qian Zhu 0003, Ying Ding 0001, David J. Wild 0001 |
BMC Bioinform. | 5 |