Chris Bizon

dblp:42/8886 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0002-9491-7674ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Explainable Enrichment-Driven GrAph Reasoner (EDGAR) for Large Knowledge Graphs with Applications in Drug Repurposing
abstract
Knowledge graphs (KGs) represent the connections and relationships between real-world entities. We propose a link prediction framework on KGs named Enrichment-Driven GrAph Reasoner (EDGAR) that infers new edges by mining entity-local rules. This approach is based on enrichment analysis, a well-established statistical method used to calculate mechanisms common to a set of differentially expressed genes. EDGAR’s inference results are inherently explainable and rankable, equipped with p-values for statistical significance of each enrichment-based rule. We demonstrate its effectiveness on a large-scale biomedical KG, ROBOKOP, focusing on drug repurposing for Alzheimer disease (AD) as a case study. Initially, we extracted 14 known drugs from the KG and identified 20 contextual biomarkers through enrichment analysis, shedding light on functional pathways relevant to the shared efficacy of drugs for AD. Subsequently, using the top 1,000 enrichment results, our enrichment-driven system identified 1,246 additional drug candidates for AD treatment. We validated the top 10 candidates using medical literature evidence. EDGAR is deployed within ROBOKOP, along with a web user interface. This is the first work to use enrichment analysis for either large graph completion or drug repurposing.
Olawumi Olasunkanmi, Evan Morris, Yaphet Kebede, Harlin Lee, Stanley C. Ahalt, Alexander Tropsha, Chris Bizon
IEEE Big Data7
2024 ExEmPLAR (Extracting, Exploring, and Embedding Pathways Leading to Actionable Research): a user-friendly interface for knowledge graph mining
abstract
SUMMARY: Knowledge graphs are being increasingly used in biomedical research to link large amounts of heterogenous data and facilitate reasoning across diverse knowledge sources. Wider adoption and exploration of knowledge graphs in the biomedical research community is limited by requirements to understand the underlying graph structure in terms of entity types and relationships, represented as nodes and edges, respectively, and learn specialized query languages for graph mining and exploration. We have developed a user-friendly interface dubbed ExEmPLAR (Extracting, Exploring, and Embedding Pathways Leading to Actionable Research) to aid reasoning over biomedical knowledge graphs and assist with data-driven research and hypothesis generation. We explain the key functionalities of ExEmPLAR and demonstrate its use with a case study considering the relationship of Trypanosoma cruzi, the etiological agent of Chagas disease, to frequently associated cardiovascular conditions. AVAILABILITY AND IMPLEMENTATION: ExEmPLAR is freely accessible at https://www.exemplar.mml.unc.edu/. For code and instructions for the using the application, see: https://github.com/beasleyjonm/AOP-COP-Path-Extractor.
Jon-Michael Beasley, Daniel R. Korn, Nyssa N. Tucker, Erick T. M. Alves, Eugene N. Muratov, Chris Bizon, Alexander Tropsha
Bioinform.6
2022 Dug: a semantic search engine leveraging peer-reviewed knowledge to query biomedical data repositories
abstract
MOTIVATION: As the number of public data resources continues to proliferate, identifying relevant datasets across heterogenous repositories is becoming critical to answering scientific questions. To help researchers navigate this data landscape, we developed Dug: a semantic search tool for biomedical datasets utilizing evidence-based relationships from curated knowledge graphs to find relevant datasets and explain why those results are returned. RESULTS: Developed through the National Heart, Lung and Blood Institute's (NHLBI) BioData Catalyst ecosystem, Dug has indexed more than 15 911 study variables from public datasets. On a manually curated search dataset, Dug's total recall (total relevant results/total results) of 0.79 outperformed default Elasticsearch's total recall of 0.76. When using synonyms or related concepts as search queries, Dug (0.36) far outperformed Elasticsearch (0.14) in terms of total recall with no significant loss in the precision of its top results. AVAILABILITY AND IMPLEMENTATION: Dug is freely available at https://github.com/helxplatform/dug. An example Dug deployment is also available for use at https://search.biodatacatalyst.renci.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander M. Waldrop, John B. Cheadle, Kira Bradford, Alexander Preiss, Robert F. Chew, Jonathan R. Holt, Yaphet Kebede, Nathan Braswell, Virginia Hench, Andrew Crerar, Chris M. Ball, Carl Schreep, P. J. Linebaugh, Hannah Hiles, Rebecca R. Boyles, Chris Bizon, Ashok K. Krishnamurthy 0001, Steven Cox 0001
Bioinform.17
2021 AI Tool with Active Learning for Detection of Rural Roadside Safety Features
abstract
Roadway safety, especially in rural areas, is one of the most critical components in transportation planning. In collaboration with North Carolina Department of Transportation (NCDOT), UNC Highway Safety Research Center (HSRC), and DOT Volpe National Transportation Systems Center, UNC Renaissance Computing Institute (RENCI) developed a roadside feature detection solution leveraging multiple convolutional neural networks. The solution used an iterative active learning (AL) computer vision model training pipeline integrated into an AI tool to detect safety features such as guardrails and utility poles in geographically distributed NC rural roads. We utilized transfer learning by adopting the Xception neural network architecture [1] as the feature extraction backbone which was then used in an iterative AL process supported by a web-based annotation tool. The annotation tool not only allowed for the collection of annotations through an iterative AL process for multiple safety features, it also enabled visual analysis and assessment of model prediction performance in the geospatial context. AL techniques were used to direct human annotators to label images that would most effectively improve the model aimed at minimizing the number of required training labels while maximizing the model’s performance. The iterative AL process combined with a common feature extraction backbone allowed fast model inference on millions of images in the AL sampling space. This enabled a rapid transition between AL rounds while also reducing the computing requirements for each round. Model feature extraction weights were then fine-tuned in the last round of AL to obtain the best accuracy. Since only about 2.7% of 2.6 million unlabeled images in the AL sampling space contain guardrails, there is a significant class imbalance problem that must be addressed in our AL sampling strategies for the guardrail classification model. In this paper, we present our AI tool processing pipeline and methodology and discuss our AL results and future work. Our AI tool can be used to detect roadside safety features and be extended to also locate them for assessing roadside hazards.
Chris Bizon, David Borland, Matthew Satusky, Robert Rittmuller, Randa Radwan, Ashok K. Krishnamurthy 0001
IEEE BigData2
2021 COVID-KOP: integrating emerging COVID-19 data with the ROBOKOP database
abstract
SUMMARY: In response to the COVID-19 pandemic, we established COVID-KOP, a new knowledgebase integrating the existing Reasoning Over Biomedical Objects linked in Knowledge Oriented Pathways (ROBOKOP) biomedical knowledge graph with information from recent biomedical literature on COVID-19 annotated in the CORD-19 collection. COVID-KOP can be used effectively to generate new hypotheses concerning repurposing of known drugs and clinical drug candidates against COVID-19 by establishing respective confirmatory pathways of drug action. AVAILABILITY AND IMPLEMENTATION: COVID-KOP is freely accessible at https://covidkop.renci.org/. For code and instructions for the original ROBOKOP, see: https://github.com/NCATS-Gamma/robokop.
Daniel R. Korn, Tesia M. Bobrowski, Michael Li, Yaphet Kebede, Patrick Wang 0003, Phillips Owen, Gaurav Vaidya, Eugene N. Muratov, Rada Chirkova, Chris Bizon, Alexander Tropsha
Bioinform.10
2021 Pre-capture multiplexing provides additional power to detect copy number variation in exome sequencing
abstract
BACKGROUND: As exome sequencing (ES) integrates into clinical practice, we should make every effort to utilize all information generated. Copy-number variation can lead to Mendelian disorders, but small copy-number variants (CNVs) often get overlooked or obscured by under-powered data collection. Many groups have developed methodology for detecting CNVs from ES, but existing methods often perform poorly for small CNVs and rely on large numbers of samples not always available to clinical laboratories. Furthermore, methods often rely on Bayesian approaches requiring user-defined priors in the setting of insufficient prior knowledge. This report first demonstrates the benefit of multiplexed exome capture (pooling samples prior to capture), then presents a novel detection algorithm, mcCNV ("multiplexed capture CNV"), built around multiplexed capture. RESULTS: We demonstrate: (1) multiplexed capture reduces inter-sample variance; (2) our mcCNV method, a novel depth-based algorithm for detecting CNVs from multiplexed capture ES data, improves the detection of small CNVs. We contrast our novel approach, agnostic to prior information, with the the commonly-used ExomeDepth. In a simulation study mcCNV demonstrated a favorable false discovery rate (FDR). When compared to calls made from matched genome sequencing, we find the mcCNV algorithm performs comparably to ExomeDepth. CONCLUSION: Implementing multiplexed capture increases power to detect single-exon CNVs. The novel mcCNV algorithm may provide a more favorable FDR than ExomeDepth. The greatest benefits of our approach derive from (1) not requiring a database of reference samples and (2) not requiring prior information about the prevalance or size of variants.
Dayne L. Filer, Fengshen Kuo, Alicia T. Brandt, Christian R. Tilley, Piotr A. Mieczkowski, Jonathan S. Berg, Kimberly Robasky, Chris Bizon, Jeffery L. Tilson, Bradford C. Powell, Darius M. Bost, Clark D. Jeffries, Kirk C. Wilhelmsen
BMC Bioinform.9
2020 Use of the Open ROBOKOP Knowledge Graph-Based Application to Provide Mechanistic Explanations for Observed Associations between Environmental Exposures and Immune-Mediated Diseases
Karamarie Fecho, Chris Bizon, Frederick W. Miller, Shepherd Schurman, Charles Schmitt, William Xue, Patrick Wang 0003, Kenneth Morton, Steven Cox 0001, Alexander Tropsha
AMIA2
2019 ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering
abstract
SUMMARY: Knowledge graphs (KGs) are quickly becoming a common-place tool for storing relationships between entities from which higher-level reasoning can be conducted. KGs are typically stored in a graph-database format, and graph-database queries can be used to answer questions of interest that have been posed by users such as biomedical researchers. For simple queries, the inclusion of direct connections in the KG and the storage and analysis of query results are straightforward; however, for complex queries, these capabilities become exponentially more challenging with each increase in complexity of the query. For instance, one relatively complex query can yield a KG with hundreds of thousands of query results. Thus, the ability to efficiently query, store, rank and explore sub-graphs of a complex KG represents a major challenge to any effort designed to exploit the use of KGs for applications in biomedical research and other domains. We present Reasoning Over Biomedical Objects linked in Knowledge Oriented Pathways as an abstraction layer and user interface to more easily query KGs and store, rank and explore query results. AVAILABILITY AND IMPLEMENTATION: An instance of the ROBOKOP UI for exploration of the ROBOKOP Knowledge Graph can be found at http://robokop.renci.org. The ROBOKOP Knowledge Graph can be accessed at http://robokopkg.renci.org. Code and instructions for building and deploying ROBOKOP are available under the MIT open software license from https://github.com/NCATS-Gamma/robokop. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kenneth Morton, Patrick Wang 0003, Chris Bizon, Steven Cox 0001, James P. Balhoff, Yaphet Kebede, Karamarie Fecho, Alexander Tropsha
Bioinform.3
2019 A novel approach for exposing and sharing clinical data: the Translator Integrated Clinical and Environmental Exposures Service
abstract
OBJECTIVE: This study aimed to develop a novel, regulatory-compliant approach for openly exposing integrated clinical and environmental exposures data: the Integrated Clinical and Environmental Exposures Service (ICEES). MATERIALS AND METHODS: The driving clinical use case for research and development of ICEES was asthma, which is a common disease influenced by hundreds of genes and a plethora of environmental exposures, including exposures to airborne pollutants. We developed a pipeline for integrating clinical data on patients with asthma-like conditions with data on environmental exposures derived from multiple public data sources. The data were integrated at the patient and visit level and used to create de-identified, binned, "integrated feature tables," which were then placed behind an OpenAPI. RESULTS: Our preliminary evaluation results demonstrate a relationship between exposure to high levels of particulate matter ≤2.5 µm in diameter (PM2.5) and the frequency of emergency department or inpatient visits for respiratory issues. For example, 16.73% of patients with average daily exposure to PM2.5 >9.62 µg/m3 experienced 2 or more emergency department or inpatient visits for respiratory issues in year 2010 compared with 7.93% of patients with lower exposures (n = 23 093). DISCUSSION: The results validated our overall approach for openly exposing and sharing integrated clinical and environmental exposures data. We plan to iteratively refine and expand ICEES by including additional years of data, feature variables, and disease cohorts. CONCLUSIONS: We believe that ICEES will serve as a regulatory-compliant model and approach for promoting open access to and sharing of integrated clinical and environmental exposures data.
Karamarie Fecho, Emily R. Pfaff, Hao Xu 0006, James Champion, Steven Cox 0001, Lisa Stillwell, David B. Peden, Chris Bizon, Ashok K. Krishnamurthy 0001, Alexander Tropsha, Stanley C. Ahalt
J. Am. Medical Informatics Assoc.8
2013 Imputation of coding variants in African Americans: better performance using data from the exome sequencing project
abstract
SUMMARY: Although the 1000 Genomes haplotypes are the most commonly used reference panel for imputation, medical sequencing projects are generating large alternate sets of sequenced samples. Imputation in African Americans using 3384 haplotypes from the Exome Sequencing Project, compared with 2184 haplotypes from 1000 Genomes Project, increased effective sample size by 8.3-11.4% for coding variants with minor allele frequency <1%. No loss of imputation quality was observed using a panel built from phenotypic extremes. We recommend using haplotypes from Exome Sequencing Project alone or concatenation of the two panels over quality score-based post-imputation selection or IMPUTE2's two-panel combination. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qing Duan, Eric Yi Liu, Paul L. Auer, Ethan M. Lange, Goo Jun, Chris Bizon, Shuo Jiao, Steven Buyske, Nora Franceschini, Chris S. Carlson, Li Hsu, Alex P. Reiner, Ulrike Peters, Jeffrey Haessler, Keith Curtis, Christina L. Wassel, Jennifer G. Robinson, Lisa W. Martin, Christopher A. Haiman, Loic Le Marchand, Tara Cox Matise, Lucia Hindorff, Dana C. Crawford, Themistocles L. Assimes, Hyun Min Kang, Gerardo Heiss, Rebecca D. Jackson, Charles L. Kooperberg, James G. Wilson, Gonçalo R. Abecasis, Kari E. North, Deborah A. Nickerson, Leslie Lange
Bioinform.7
2012 ReQON: a Bioconductor package for recalibrating quality scores from next-generation sequencing data
abstract
BACKGROUND: Next-generation sequencing technologies have become important tools for genome-wide studies. However, the quality scores that are assigned to each base have been shown to be inaccurate. If the quality scores are used in downstream analyses, these inaccuracies can have a significant impact on the results. RESULTS: Here we present ReQON, a tool that recalibrates the base quality scores from an input BAM file of aligned sequencing data using logistic regression. ReQON also generates diagnostic plots showing the effectiveness of the recalibration. We show that ReQON produces quality scores that are both more accurate, in the sense that they more closely correspond to the probability of a sequencing error, and do a better job of discriminating between sequencing errors and non-errors than the original quality scores. We also compare ReQON to other available recalibration tools and show that ReQON is less biased and performs favorably in terms of quality score accuracy. CONCLUSION: ReQON is an open source software package, written in R and available through Bioconductor, for recalibrating base quality scores for next-generation sequencing data. ReQON produces a new BAM file with more accurate quality scores, which can improve the results of downstream analysis, and produces several diagnostic plots showing the effectiveness of the recalibration.
Christopher R. Cabanski, Keary Cavin, Chris Bizon, Matthew D. Wilkerson, Joel S. Parker, Kirk C. Wilhelmsen, Charles M. Perou, J. S. Marron, D. Neil Hayes
BMC Bioinform.3
2012 VisualDecisionLinc: A visual analytics approach for comparative effectiveness-based clinical decision support in psychiatry
Ketan K. Mane, Chris Bizon, Charles Schmitt, Phillips Owen, Bruce Burchett, Ricardo Pietrobon, Kenneth Gersing
J. Biomed. Informatics2
2010 Pre-calculated protein structure alignments at the RCSB PDB website
abstract
SUMMARY: With the continuous growth of the RCSB Protein Data Bank (PDB), providing an up-to-date systematic structure comparison of all protein structures poses an ever growing challenge. Here, we present a comparison tool for calculating both 1D protein sequence and 3D protein structure alignments. This tool supports various applications at the RCSB PDB website. First, a structure alignment web service calculates pairwise alignments. Second, a stand-alone application runs alignments locally and visualizes the results. Third, pre-calculated 3D structure comparisons for the whole PDB are provided and updated on a weekly basis. These three applications allow users to discover novel relationships between proteins available either at the RCSB PDB or provided by the user. AVAILABILITY AND IMPLEMENTATION: A web user interface is available at http://www.rcsb.org/pdb/workbench/workbench.do. The source code is available under the LGPL license from http://www.biojava.org. A source bundle, prepared for local execution, is available from http://source.rcsb.org CONTACT: [email protected]; [email protected].
Andreas Prlic, Spencer Bliven, Peter W. Rose, Wolfgang Bluhm, Chris Bizon, Adam Godzik, Philip E. Bourne
Bioinform.5