Richard A. Gibbs

dblp:88/2875 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0002-1356-5698ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics › structural variation
structural variant genotyping
0.512021
muCNV: genotyping structural variants for population-level sequencing · Bioinform. 2021
Bioinformatics and computational biology › genomics
population-scale sequencing
0.112021
muCNV: genotyping structural variants for population-level sequencing · Bioinform. 2021
Bioinformatics and computational biology › genomics › genomic data management
genome database
0.012000
The Human Transcript Database: a catalogue of full length cDNA inserts · Bioinform. 2000
Bioinformatics and computational biology
gel electrophoresis
0.011995
Accurate determination of DNA in agarose gels using the novel algorithm GelScann(1.0) · Comput. Appl. Biosci. 1995
Bioinformatics and computational biology › bioinformatics infrastructure
molecular biology software
0.011995
Accurate determination of DNA in agarose gels using the novel algorithm GelScann(1.0) · Comput. Appl. Biosci. 1995

Methods — techniques the papers use, named apart from their topics

pileup aggregation · 0.5statistical discrimination · 0.0
YearPublicationVenuePosition
2024 Empowering personalized pharmacogenomics with generative AI solutions
abstract
OBJECTIVE: This study evaluates an AI assistant developed using OpenAI's GPT-4 for interpreting pharmacogenomic (PGx) testing results, aiming to improve decision-making and knowledge sharing in clinical genetics and to enhance patient care with equitable access. MATERIALS AND METHODS: The AI assistant employs retrieval-augmented generation (RAG), which combines retrieval and generative techniques, by harnessing a knowledge base (KB) that comprises data from the Clinical Pharmacogenetics Implementation Consortium (CPIC). It uses context-aware GPT-4 to generate tailored responses to user queries from this KB, further refined through prompt engineering and guardrails. RESULTS: Evaluated against a specialized PGx question catalog, the AI assistant showed high efficacy in addressing user queries. Compared with OpenAI's ChatGPT 3.5, it demonstrated better performance, especially in provider-specific queries requiring specialized data and citations. Key areas for improvement include enhancing accuracy, relevancy, and representative language in responses. DISCUSSION: The integration of context-aware GPT-4 with RAG significantly enhanced the AI assistant's utility. RAG's ability to incorporate domain-specific CPIC data, including recent literature, proved beneficial. Challenges persist, such as the need for specialized genetic/PGx models to improve accuracy and relevancy and addressing ethical, regulatory, and safety concerns. CONCLUSION: This study underscores generative AI's potential for transforming healthcare provider support and patient accessibility to complex pharmacogenomic information. While careful implementation of large language models like GPT-4 is necessary, it is clear that they can substantially improve understanding of pharmacogenomic data. With further development, these tools could augment healthcare expertise, provider productivity, and the delivery of equitable, patient-centered healthcare services.
Mullai Murugan, Eric Venner, Christie M. Ballantyne, Katherine M. Robinson, James C. Coons, Philip E. Empey, Richard A. Gibbs
J. Am. Medical Informatics Assoc.9
2021 muCNV: genotyping structural variants for population-level sequencing
abstract
MOTIVATION: There are high demands for joint genotyping of structural variations with short-read sequencing, but efficient and accurate genotyping in population scale is a challenging task. RESULTS: We developed muCNV that aggregates per-sample summary pileups for joint genotyping of > 100,000 samples. Pilot results show very low Mendelian inconsistencies. Applications to large-scale projects in cloud show the computational efficiencies of muCNV genotyping pipeline. AVAILABILITY: muCNV is publicly available for download at: https://github.com/gjun/muCNV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Goo Jun, Fritz J. Sedlazeck, Qihui Zhu, Adam English, Ginger Metcalf, Hyun Min Kang, Charles Lee, Richard A. Gibbs, Eric Boerwinkle
Bioinform.8
2021 Genomic considerations for FHIR®; eMERGE implementation lessons
Mullai Murugan, Lawrence J. Babb, Casey Overby Taylor, Luke V. Rasmussen, Robert R. Freimuth, Eric Venner, Victoria Yi, Stephen Granite, Hana Zouk, Samuel J. Aronson, Kevin Power, Alexander Fedotov, David R. Crosslin, David Fasel, Gail P. Jarvik, Hakon Hakonarson, Hana Bangash, Iftikhar J. Kullo, John J. Connolly, Jordan G. Nestor, Pedro J. Caraballo, Wei-Qi Wei, Ken Wiley, Heidi L. Rehm, Richard A. Gibbs
J. Biomed. Informatics26
2019 ARBoR: an identity and security solution for clinical reporting
abstract
MOTIVATION: Clinical genome sequencing laboratories return reports containing clinical testing results, signed by a board-certified clinical geneticist, to the ordering physician. This report is often a PDF, but can also be a paper copy or a structured data file. The reports are frequently modified and reissued due to changes in variant interpretation or clinical attributes. MATERIALS AND METHODS: To precisely track report authenticity, we developed ARBoR (Authenticated Resources in a Hashed Block Registry), an application for tracking the authenticity and lineage of versioned clinical reports even when they are distributed as PDF or paper copies. ARBoR tracks clinical reports as cryptographically signed hash blocks in an electronic ledger file, which is then exactly replicated to many clients. RESULTS: ARBoR was implemented for clinical reporting in the Human Genome Sequencing Center Clinical Laboratory, initially as part of the National Institute of Health's Electronic Medical Record and Genomics (eMERGE) project. CONCLUSIONS: To date, we have issued 15 205 versioned clinical reports tracked by ARBoR. This system has provided us with a simple and tamper-proof mechanism for tracking clinical reports with a complicated update history.
Eric Venner, Mullai Murugan, Walker Hale, Jordan M. Jones, Victoria Yi, Richard A. Gibbs
J. Am. Medical Informatics Assoc.7
2018 Empowering genomic medicine by establishing critical sequencing result data flows: the eMERGE example
abstract
The eMERGE Network is establishing methods for electronic transmittal of patient genetic test results from laboratories to healthcare providers across organizational boundaries. We surveyed the capabilities and needs of different network participants, established a common transfer format, and implemented transfer mechanisms based on this format. The interfaces we created are examples of the connectivity that must be instantiated before electronic genetic and genomic clinical decision support can be effectively built at the point of care. This work serves as a case example for both standards bodies and other organizations working to build the infrastructure required to provide better electronic clinical decision support for clinicians.
Samuel J. Aronson, Lawrence J. Babb, Darren C. Ames, Richard A. Gibbs, Eric Venner, John J. Connelly, Keith Marsolo, Chunhua Weng, Marc S. Williams, Andrea L. Hartzler, Wayne H. Liang, James D. Ralston, Emily Beth Devine, Shawn N. Murphy, Christopher G. Chute, Pedro J. Caraballo, Iftikhar J. Kullo, Robert R. Freimuth, Luke V. Rasmussen, Firas H. Wehbe, Josh F. Peterson, Jamie R. Robinson, Ken Wiley, Casey Overby Taylor
J. Am. Medical Informatics Assoc.4
2017 SeqCNV: a novel method for identification of copy number variations in targeted next-generation sequencing data
abstract
BACKGROUND: Targeted next-generation sequencing (NGS) has been widely used as a cost-effective way to identify the genetic basis of human disorders. Copy number variations (CNVs) contribute significantly to human genomic variability, some of which can lead to disease. However, effective detection of CNVs from targeted capture sequencing data remains challenging. RESULTS: Here we present SeqCNV, a novel CNV calling method designed to use capture NGS data. SeqCNV extracts the read depth information and utilizes the maximum penalized likelihood estimation (MPLE) model to identify the copy number ratio and CNV boundary. We applied SeqCNV to both bacterial artificial clone (BAC) and human patient NGS data to identify CNVs. These CNVs were validated by array comparative genomic hybridization (aCGH). CONCLUSIONS: SeqCNV is able to robustly identify CNVs of different size using capture NGS data. Compared with other CNV-calling methods, SeqCNV shows a significant improvement in both sensitivity and specificity.
Yong Chen 0016, Ming Cao 0005, Violet Gelowani, Mingchu Xu, Smriti A. Agrawal, Yumei Li 0007, Stephen P. Daiger, Richard A. Gibbs, Fei Wang 0017, Rui Chen 0013
BMC Bioinform.10
2016 DNAism: exploring genomic datasets on the web with Horizon Charts
abstract
BACKGROUND: Computational biologists daily face the need to explore massive amounts of genomic data. New visualization techniques can help researchers navigate and understand these big data. Horizon Charts are a relatively new visualization method that, under the right circumstances, maximizes data density without losing graphical perception. RESULTS: Horizon Charts have been successfully applied to understand multi-metric time series data. We have adapted an existing JavaScript library (Cubism) that implements Horizon Charts for the time series domain so that it works effectively with genomic datasets. We call this new library DNAism. CONCLUSIONS: Horizon Charts can be an effective visual tool to explore complex and large genomic datasets. Researchers can use our library to leverage these techniques to extract additional insights from their own datasets.
David Rio Deiros, Richard A. Gibbs, Jeffrey Rogers
BMC Bioinform.2
2016 A hybrid computational strategy to address WGS variant analysis in >5000 samples
abstract
BACKGROUND: The decreasing costs of sequencing are driving the need for cost effective and real time variant calling of whole genome sequencing data. The scale of these projects are far beyond the capacity of typical computing resources available with most research labs. Other infrastructures like the cloud AWS environment and supercomputers also have limitations due to which large scale joint variant calling becomes infeasible, and infrastructure specific variant calling strategies either fail to scale up to large datasets or abandon joint calling strategies. RESULTS: We present a high throughput framework including multiple variant callers for single nucleotide variant (SNV) calling, which leverages hybrid computing infrastructure consisting of cloud AWS, supercomputers and local high performance computing infrastructures. We present a novel binning approach for large scale joint variant calling and imputation which can scale up to over 10,000 samples while producing SNV callsets with high sensitivity and specificity. As a proof of principle, we present results of analysis on Cohorts for Heart And Aging Research in Genomic Epidemiology (CHARGE) WGS freeze 3 dataset in which joint calling, imputation and phasing of over 5300 whole genome samples was produced in under 6 weeks using four state-of-the-art callers. The callers used were SNPTools, GATK-HaplotypeCaller, GATK-UnifiedGenotyper and GotCloud. We used Amazon AWS, a 4000-core in-house cluster at Baylor College of Medicine, IBM power PC Blue BioU at Rice and Rhea at Oak Ridge National Laboratory (ORNL) for the computation. AWS was used for joint calling of 180 TB of BAM files, and ORNL and Rice supercomputers were used for the imputation and phasing step. All other steps were carried out on the local compute cluster. The entire operation used 5.2 million core hours and only transferred a total of 6 TB of data across the platforms. CONCLUSIONS: Even with increasing sizes of whole genome datasets, ensemble joint calling of SNVs for low coverage data can be accomplished in a scalable, cost effective and fast manner by using heterogeneous computing platforms without compromising on the quality of variants.
Zhuoyi Huang, Navin Rustagi, Narayanan Veeraraghavan, Andrew Carroll, Richard A. Gibbs, Eric Boerwinkle, Manjunath Gorentla Venkata, Fuli Yu
BMC Bioinform.5
2016 ITD assembler: an algorithm for internal tandem duplication discovery from short-read sequencing data
abstract
BACKGROUND: Detection of tandem duplication within coding exons, referred to as internal tandem duplication (ITD), remains challenging due to inefficiencies in alignment of ITD-containing reads to the reference genome. There is a critical need to develop efficient methods to recover these important mutational events. RESULTS: In this paper we introduce ITD Assembler, a novel approach that rapidly evaluates all unmapped and partially mapped reads from whole exome NGS data using a De Bruijn graphs approach to select reads that harbor cycles of appropriate length, followed by assembly using overlap-layout-consensus. We tested ITD Assembler on The Cancer Genome Atlas AML dataset as a truth set. ITD Assembler identified the highest percentage of reported FLT3-ITDs when compared to other ITD detection algorithms, and discovered additional ITDs in FLT3, KIT, CEBPA, WT1 and other genes. Evidence of polymorphic ITDs in 54 genes were also found. Novel ITDs were validated by analyzing the corresponding RNA sequencing data. CONCLUSIONS: ITD Assembler is a very sensitive tool which can detect partial, large and complex tandem duplications. This study highlights the need to more effectively look for ITD's in other cancers and Mendelian diseases.
Navin Rustagi, Oliver A. Hampton, Liu Xi, Richard A. Gibbs, Sharon E. Plon, Marek Kimmel, David A. Wheeler
BMC Bioinform.5
2014 Launching genomics into the cloud: deployment of Mercury, a next generation sequence analysis pipeline
abstract
BACKGROUND: Massively parallel DNA sequencing generates staggering amounts of data. Decreasing cost, increasing throughput, and improved annotation have expanded the diversity of genomics applications in research and clinical practice. This expanding scale creates analytical challenges: accommodating peak compute demand, coordinating secure access for multiple analysts, and sharing validated tools and results. RESULTS: To address these challenges, we have developed the Mercury analysis pipeline and deployed it in local hardware and the Amazon Web Services cloud via the DNAnexus platform. Mercury is an automated, flexible, and extensible analysis workflow that provides accurate and reproducible genomic results at scales ranging from individuals to large cohorts. CONCLUSIONS: By taking advantage of cloud computing and with Mercury implemented on the DNAnexus platform, we have demonstrated a powerful combination of a robust and fully validated software pipeline and a scalable computational resource that, to date, we have applied to more than 10,000 whole genome and whole exome samples.
Jeffrey G. Reid, Andrew Carroll, Narayanan Veeraraghavan, Mahmoud Dahdouli, Andreas Sundquist, Adam C. English, Matthew N. Bainbridge, Simon White, William J. Salerno, Christian Buhay, Fuli Yu, Donna M. Muzny, Richard Daly, Geoff Duyk, Richard A. Gibbs, Eric Boerwinkle
BMC Bioinform.15
2012 An integrative variant analysis suite for whole exome next-generation sequencing data
abstract
BACKGROUND: Whole exome capture sequencing allows researchers to cost-effectively sequence the coding regions of the genome. Although the exome capture sequencing methods have become routine and well established, there is currently a lack of tools specialized for variant calling in this type of data. RESULTS: Using statistical models trained on validated whole-exome capture sequencing data, the Atlas2 Suite is an integrative variant analysis pipeline optimized for variant discovery on all three of the widely used next generation sequencing platforms (SOLiD, Illumina, and Roche 454). The suite employs logistic regression models in conjunction with user-adjustable cutoffs to accurately separate true SNPs and INDELs from sequencing and mapping errors with high sensitivity (96.7%). CONCLUSION: We have implemented the Atlas2 Suite and applied it to 92 whole exome samples from the 1000 Genomes Project. The Atlas2 Suite is available for download at http://sourceforge.net/projects/atlas2/. In addition to a command line version, the suite has been integrated into the Genboree Workbench, allowing biomedical scientists with minimal informatics expertise to remotely call, view, and further analyze variants through a simple web interface. The existing genomic databases displayed via the Genboree browser also streamline the process from variant discovery to functional genomics analysis, resulting in an off-the-shelf toolkit for the broader community.
Danny Challis, Jin Yu 0003, Uday S. Evani, Andrew R. Jackson, Sameer Paithankar, Cristian Coarfa, Aleksandar Milosavljevic, Richard A. Gibbs, Fuli Yu
BMC Bioinform.8
2005 SNPdetector: A Software Tool for Sensitive and Accurate SNP Detection
abstract
Identification of single nucleotide polymorphisms (SNPs) and mutations is important for the discovery of genetic predisposition to complex diseases. PCR resequencing is the method of choice for de novo SNP discovery. However, manual curation of putative SNPs has been a major bottleneck in the application of this method to high-throughput screening. Therefore it is critical to develop a more sensitive and accurate computational method for automated SNP detection. We developed a software tool, SNPdetector, for automated identification of SNPs and mutations in fluorescence-based resequencing reads. SNPdetector was designed to model the process of human visual inspection and has a very low false positive and false negative rate. We demonstrate the superior performance of SNPdetector in SNP and mutation analysis by comparing its results with those derived by human inspection, PolyPhred (a popular SNP detection tool), and independent genotype assays in three large-scale investigations. The first study identified and validated inter- and intra-subspecies variations in 4,650 traces of 25 inbred mouse strains that belong to either the Mus musculus species or the M. spretus species. Unexpected heterozygosity in CAST/Ei strain was observed in two out of 1,167 mouse SNPs. The second study identified 11,241 candidate SNPs in five ENCODE regions of the human genome covering 2.5 Mb of genomic sequence. Approximately 50% of the candidate SNPs were selected for experimental genotyping; the validation rate exceeded 95%. The third study detected ENU-induced mutations (at 0.04% allele frequency) in 64,896 traces of 1,236 zebra fish. Our analysis of three large and diverse test datasets demonstrated that SNPdetector is an effective tool for genome-scale research and for large-sample clinical studies. SNPdetector runs on Unix/Linux platform and is available publicly (http://lpg.nci.nih.gov).
David A. Wheeler, Imtiaz Yakub, Sharon Wei, Raman Sood, William Rowe, Paul P. Liu, Richard A. Gibbs, Kenneth H. Buetow
PLoS Comput. Biol.8
2000 The Human Transcript Database: a catalogue of full length cDNA inserts
abstract
SUMMARY: Full length cDNA sequences are an important resource for the research community but are currently intermingled with other sequences. We have identified the human full length insert cDNA sequences in GenBank and placed them in a single location, the Human Transcript Database. AVAILIBILITY: The Human Transcript Database is available at http://www.hgsc.bcm.tms.edu/HTDB/. CONTACT: John Bouck: [email protected]
John Bouck, Michael P. McLeod, Kim Worley, Richard A. Gibbs
Bioinform.4
1995 Accurate determination of DNA in agarose gels using the novel algorithm GelScann(1.0)
abstract
GelScann(1.0) is a user-friendly program that accurately quantitates DNA from CCD imaged agarose gels. The algorithm automatically locates lanes, locates bands within a given lane, and quantitates the intensity of each band. GelScann (1.0) uses a statistical method to discriminate DNA from local background fluorescence. The sensitivity of the LaneFinder and BandFinder can be adjusted by the user interacting with GelScann (1.0)'s option box. We demonstrate that GelScann(1.0) can accurately determine DNA that GelScann(1.0) can accurately determine DNA concentrations for routine molecular biology applications and determine the relative intensities of PCR amplified fragments for genetic testing.
Michael L. Metzker, Kyle M. Allain, Richard A. Gibbs
Comput. Appl. Biosci.3