Gilbert S. Omenn

dblp:35/7105 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-8976-6074ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 6 since 2021
YearPublicationVenuePosition
2025 AI-Based Computational Methods in Early Drug Discovery and Post Market Drug Assessment: A Survey
abstract
Over the past few years, artificial intelligence (AI) has emerged as a transformative force in drug discovery and development (DDD), revolutionizing many aspects of the process. This survey provides a comprehensive review of recent advancements in AI applications within early drug discovery and post-market drug assessment. It addresses the identification and prioritization of new therapeutic targets, prediction of drug-target interaction (DTI), design of novel drug-like molecules, and assessment of the clinical efficacy of new medications. By integrating AI technologies, pharmaceutical companies can accelerate the discovery of new treatments, enhance the precision of drug development, and bring more effective therapies to market. This shift represents a significant move towards more efficient and cost-effective methodologies in the DDD landscape.
Flora Rajaei, Cristian Minoccheri, Emily Wittrup, Richard C. Wilson 0001, Brian D. Athey, Gilbert S. Omenn, Kayvan Najarian
IEEE Trans. Comput. Biol. Bioinform.6
2023 Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes
J. Biomed. Informatics26
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics37
2022 Fast and accurate Ab Initio Protein structure prediction using deep learning potentials
abstract
Despite the immense progress recently witnessed in protein structure prediction, the modeling accuracy for proteins that lack sequence and/or structure homologs remains to be improved. We developed an open-source program, DeepFold, which integrates spatial restraints predicted by multi-task deep residual neural-networks along with a knowledge-based energy function to guide its gradient-descent folding simulations. The results on large-scale benchmark tests showed that DeepFold creates full-length models with accuracy significantly beyond classical folding approaches and other leading deep learning methods. Of particular interest is the modeling performance on the most difficult targets with very few homologous sequences, where DeepFold achieved an average TM-score that was 40.3% higher than trRosetta and 44.9% higher than DMPfold. Furthermore, the folding simulations for DeepFold were 262 times faster than traditional fragment assembly simulations. These results demonstrate the power of accurately predicted deep learning potentials to improve both the accuracy and speed of ab initio protein structure prediction.
Robin Pearce, Yang Li 0107, Gilbert S. Omenn, Yang Zhang 0040
PLoS Comput. Biol.3
2021 IsoResolve: predicting splice isoform functions by integrating gene and isoform-level features with domain adaptation
abstract
MOTIVATION: High resolution annotation of gene functions is a central goal in functional genomics. A single gene may produce multiple isoforms with different functions through alternative splicing. Conventional approaches, however, consider a gene as a single entity without differentiating these functionally different isoforms. Towards understanding gene functions at higher resolution, recent efforts have focused on predicting the functions of isoforms. However, the performance of existing methods is far from satisfactory mainly because of the lack of isoform-level functional annotation. RESULTS: We present IsoResolve, a novel approach for isoform function prediction, which leverages the information from gene function prediction models with domain adaptation (DA). IsoResolve treats gene-level and isoform-level features as source and target domains, respectively. It uses DA to project the two domains into a latent variable space in such a way that the latent variables from the two domains have similar distribution, which enables the gene domain information to be leveraged for isoform function prediction. We systematically evaluated the performance of IsoResolve in predicting functions. Compared with five state-of-the-art methods, IsoResolve achieved significantly better performance. IsoResolve was further validated by case studies of genes with isoform-level functional annotation. AVAILABILITY AND IMPLEMENTATION: IsoResolve is freely available at https://github.com/genemine/IsoResolve. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hong-Dong Li, Changhuo Yang, Mengyun Yang, Fang-Xiang Wu, Gilbert S. Omenn, Jianxin Wang 0001
Bioinform.6
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.17
2018 Prioritizing predictive biomarkers for gene essentiality in cancer cells with mRNA expression data and DNA copy number profile
abstract
Motivation: Finding driver genes that are responsible for the aberrant proliferation rate of cancer cells is informative for both cancer research and the development of targeted drugs. The established experimental and computational methods are labor-intensive. To make algorithms feasible in real clinical settings, methods that can predict driver genes using less experimental data are urgently needed. Results: We designed an effective feature selection method and used Support Vector Machines (SVM) to predict the essentiality of the potential driver genes in cancer cell lines with only 10 genes as features. The accuracy of our predictions was the highest in the Broad-DREAM Gene Essentiality Prediction Challenge. We also found a set of genes whose essentiality could be predicted much more accurately than others, which we called Accurately Predicted (AP) genes. Our method can serve as a new way of assessing the essentiality of genes in cancer cells. Availability and implementation: The raw data that support the findings of this study are available at Synapse. https://www.synapse.org/#! Synapse: syn2384331/wiki/62825. Source code is available at GitHub. https://github.com/GuanLab/DREAM-Gene-Essentiality-Challenge. Supplementary information: Supplementary data are available at Bioinformatics online.
Yuanfang Guan, Tingyang Li, Hongjiu Zhang, Gilbert S. Omenn
Bioinform.5
2017 miRmine: a database of human miRNA expression profiles
abstract
MOTIVATION: MicroRNAs (miRNAs) are small non-coding RNAs that are involved in post-transcriptional regulation of gene expression. In this high-throughput sequencing era, a tremendous amount of RNA-seq data is accumulating, and full utilization of publicly available miRNA data is an important challenge. These data are useful to determine expression values for each miRNA, but quantification pipelines are in a primitive stage and still evolving; there are many factors that affect expression values significantly. RESULTS: We used 304 high-quality microRNA sequencing (miRNA-seq) datasets from NCBI-SRA and calculated expression profiles for different tissues and cell-lines. In each miRNA-seq dataset, we found an average of more than 500 miRNAs with higher than 5x coverage, and we explored the top five highly expressed miRNAs in each tissue and cell-line. This user-friendly miRmine database has options to retrieve expression profiles of single or multiple miRNAs for a specific tissue or cell-line, either normal or with disease information. Results can be displayed in multiple interactive, graphical and downloadable formats. AVAILABILITY AND IMPLEMENTATION: http://guanlab.ccmb.med.umich.edu/mirmine. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bharat Panwar, Gilbert S. Omenn, Yuanfang Guan
Bioinform.2
2016 A proteogenomic approach to understand splice isoform functions through sequence and expression-based computational modeling
abstract
The products of multi-exon genes are a mixture of alternatively spliced isoforms, from which the translated proteins can have similar, different or even opposing functions. It is therefore essential to differentiate and annotate functions for individual isoforms. Computational approaches provide an efficient complement to expensive and time-consuming experimental studies. The input data of these methods range from DNA sequence, to RNA selection pressure, to expressed sequence tags, to full-length complementary DNA, to exon array, to RNA-seq expression, to proteomic data. Notably, RNA-seq technology generates quantitative profiling of transcript expression at the genome scale, with an unprecedented amount of expression data available for developing isoform function prediction methods. Integrative analysis of these data at different molecular levels enables a proteogenomic approach to systematically interrogate isoform functions. Here, we briefly review the state-of-the-art methods according to their input data sources, discuss their advantages and limitations and point out potential ways to improve prediction accuracies.
Hongdong Li, Gilbert S. Omenn, Yuanfang Guan
Briefings Bioinform.2
2015 Development of data representation standards by the human proteome organization proteomics standards initiative
abstract
OBJECTIVE: To describe the goals of the Proteomics Standards Initiative (PSI) of the Human Proteome Organization, the methods that the PSI has employed to create data standards, the resulting output of the PSI, lessons learned from the PSI's evolution, and future directions and synergies for the group. MATERIALS AND METHODS: The PSI has 5 categories of deliverables that have guided the group. These are minimum information guidelines, data formats, controlled vocabularies, resources and software tools, and dissemination activities. These deliverables are produced via the leadership and working group organization of the initiative, driven by frequent workshops and ongoing communication within the working groups. Official standards are subjected to a rigorous document process that includes several levels of peer review prior to release. RESULTS: We have produced and published minimum information guidelines describing what information should be provided when making data public, either via public repositories or other means. The PSI has produced a series of standard formats covering mass spectrometer input, mass spectrometer output, results of informatics analysis (both qualitative and quantitative analyses), reports of molecular interaction data, and gel electrophoresis analyses. We have produced controlled vocabularies that ensure that concepts are uniformly annotated in the formats and engaged in extensive software development and dissemination efforts so that the standards can efficiently be used by the community.Conclusion In its first dozen years of operation, the PSI has produced many standards that have accelerated the field of proteomics by facilitating data exchange and deposition to data repositories. We look to the future to continue developing standards for new proteomics technologies and workflows and mechanisms for integration with other omics data types. Our products facilitate the translation of genomics and proteomics findings to clinical and biological phenotypes. The PSI website can be accessed at http://www.psidev.info.
Eric W. Deutsch, Juan P. Albar, Pierre-Alain Binz, Martin Eisenacher, Andrew R. Jones, Gerhard Mayer, Gilbert S. Omenn, Sandra E. Orchard, Juan Antonio Vizcaíno, Henning Hermjakob
J. Am. Medical Informatics Assoc.7
2014 MetDisease - connecting metabolites to diseases via literature
abstract
MOTIVATION: In recent years, metabolomics has emerged as an approach to perform large-scale characterization of small molecules in biological systems. Metabolomics posed a number of bioinformatics challenges associated in data analysis and interpretation. Genome-based metabolic reconstructions have established a powerful framework for connecting metabolites to genes through metabolic reactions and enzymes that catalyze them. Pathway databases and bioinformatics tools that use this framework have proven to be useful for annotating experimental metabolomics data. This framework can be used to infer connections between metabolites and diseases through annotated disease genes. However, only about half of experimentally detected metabolites can be mapped to canonical metabolic pathways. We present a new Cytoscape 3 plug-in, MetDisease, which uses an alternative approach to link metabolites to disease information. MetDisease uses Medical Subject Headings (MeSH) disease terms mapped to PubChem compounds through literature to annotate compound networks. AVAILABILITY AND IMPLEMENTATION: MetDisease can be downloaded from http://apps.cytoscape.org/apps/metdisease or installed via the Cytoscape app manager. Further information about MetDisease can be found at http://metdisease.ncibi.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary Data are available at Bioinformatics online.
William Duren, Terry E. Weymouth, Tim Hull, Gilbert S. Omenn, Brian D. Athey, Charles F. Burant, Alla Karnovsky
Bioinform.4
2013 Healthcare information technology and economics
abstract
At the 2011 American College of Medical Informatics (ACMI) Winter Symposium we studied the overlap between health IT and economics and what leading healthcare delivery organizations are achieving today using IT that might offer paths for the nation to follow for using health IT in healthcare reform. We recognized that health IT by itself can improve health value, but its main contribution to health value may be that it can make possible new care delivery models to achieve much larger value. Health IT is a critically important enabler to fundamental healthcare system changes that may be a way out of our current, severe problem of rising costs and national deficit. We review the current state of healthcare costs, federal health IT stimulus programs, and experiences of several leading organizations, and offer a model for how health IT fits into our health economic future.
Thomas H. Payne, David W. Bates, Eta S. Berner, Elmer V. Bernstam, H. Dominic Covvey, Mark E. Frisse, Thomas Graf, Robert A. Greenes, Edward P. Hoffer, Gilad J. Kuperman, Harold P. Lehmann, Louise Liang, Blackford Middleton, Gilbert S. Omenn, Judy G. Ozbolt
J. Am. Medical Informatics Assoc.14
2013 Systematically Differentiating Functions for Alternatively Spliced Isoforms through Integrating RNA-seq Data
abstract
Integrating large-scale functional genomic data has significantly accelerated our understanding of gene functions. However, no algorithm has been developed to differentiate functions for isoforms of the same gene using high-throughput genomic data. This is because standard supervised learning requires 'ground-truth' functional annotations, which are lacking at the isoform level. To address this challenge, we developed a generic framework that interrogates public RNA-seq data at the transcript level to differentiate functions for alternatively spliced isoforms. For a specific function, our algorithm identifies the 'responsible' isoform(s) of a gene and generates classifying models at the isoform level instead of at the gene level. Through cross-validation, we demonstrated that our algorithm is effective in assigning functions to genes, especially the ones with multiple isoforms, and robust to gene expression levels and removal of homologous gene pairs. We identified genes in the mouse whose isoforms are predicted to have disparate functionalities and experimentally validated the 'responsible' isoforms using data from mammary tissue. With protein structure modeling and experimental evidence, we further validated the predicted isoform functional differences for the genes Cdkn2a and Anxa6. Our generic framework is the first to predict and differentiate functions for alternatively spliced isoforms, instead of genes, using genomic data. It is extendable to any base machine learner and other species with alternatively spliced isoforms, and shifts the current gene-centered function prediction to isoform-level predictions.
Ridvan Eksi, Hongdong Li, Rajasree Menon, Yuchen Wen, Gilbert S. Omenn, Matthias Kretzler, Yuanfang Guan
PLoS Comput. Biol.5
2012 Metscape 2 bioinformatics tool for the analysis and visualization of metabolomics and gene expression data
abstract
MOTIVATION: Metabolomics is a rapidly evolving field that holds promise to provide insights into genotype-phenotype relationships in cancers, diabetes and other complex diseases. One of the major informatics challenges is providing tools that link metabolite data with other types of high-throughput molecular data (e.g. transcriptomics, proteomics), and incorporate prior knowledge of pathways and molecular interactions. RESULTS: We describe a new, substantially redesigned version of our tool Metscape that allows users to enter experimental data for metabolites, genes and pathways and display them in the context of relevant metabolic networks. Metscape 2 uses an internal relational database that integrates data from KEGG and EHMN databases. The new version of the tool allows users to identify enriched pathways from expression profiling data, build and analyze the networks of genes and metabolites, and visualize changes in the gene/metabolite data. We demonstrate the applications of Metscape to annotate molecular pathways for human and mouse metabolites implicated in the pathogenesis of sepsis-induced acute lung injury, for the analysis of gene expression and metabolite data from pancreatic ductal adenocarcinoma, and for identification of the candidate metabolites involved in cancer and inflammation. AVAILABILITY: Metscape is part of the National Institutes of Health-supported National Center for Integrative Biomedical Informatics (NCIBI) suite of tools, freely available at http://metscape.ncibi.org. It can be downloaded from http://cytoscape.org or installed via Cytoscape plugin manager. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alla Karnovsky, Terry E. Weymouth, Tim Hull, V. Glenn Tarcea, Giovanni Scardoni, Carlo Laudanna, Maureen A. Sartor, Kathleen A. Stringer, H. V. Jagadish, Charles F. Burant, Brian D. Athey, Gilbert S. Omenn
Bioinform.12
2012 Metab2MeSH: annotating compounds with medical subject headings
abstract
SUMMARY: Progress in high-throughput genomic technologies has led to the development of a variety of resources that link genes to functional information contained in the biomedical literature. However, tools attempting to link small molecules to normal and diseased physiology and published data relevant to biologists and clinical investigators, are still lacking. With metabolomics rapidly emerging as a new omics field, the task of annotating small molecule metabolites becomes highly relevant. Our tool Metab2MeSH uses a statistical approach to reliably and automatically annotate compounds with concepts defined in Medical Subject Headings, and the National Library of Medicine's controlled vocabulary for biomedical concepts. These annotations provide links from compounds to biomedical literature and complement existing resources such as PubChem and the Human Metabolome Database.
Maureen A. Sartor, Alexander S. Ade, Zachary C. Wright, David J. States, Gilbert S. Omenn, Brian D. Athey, Alla Karnovsky
Bioinform.5
2012 The NIH National Center for Integrative Biomedical Informatics (NCIBI)
abstract
The National Center for Integrative and Biomedical Informatics (NCIBI) is one of the eight NCBCs. NCIBI supports information access and data analysis for biomedical researchers, enabling them to build computational and knowledge models of biological systems to address the Driving Biological Problems (DBPs). The NCIBI DBPs have included prostate cancer progression, organ-specific complications of type 1 and 2 diabetes, bipolar disorder, and metabolic analysis of obesity syndrome. Collaborating with these and other partners, NCIBI has developed a series of software tools for exploratory analysis, concept visualization, and literature searches, as well as core database and web services resources. Many of our training and outreach initiatives have been in collaboration with the Research Centers at Minority Institutions (RCMI), integrating NCIBI and RCMI faculty and students, culminating each year in an annual workshop. Our future directions include focusing on the TranSMART data sharing and analysis initiative.
Brian D. Athey, James D. Cavalcoli, H. V. Jagadish, Gilbert S. Omenn, Barbara Mirel, Matthias Kretzler, Charles F. Burant, Raphael D. Isokpehi, Charles DeLisi
J. Am. Medical Informatics Assoc.4
2010 Metscape: a Cytoscape plug-in for visualizing and interpreting metabolomic data in the context of human metabolic networks
abstract
SUMMARY: Metscape is a plug-in for Cytoscape, used to visualize and interpret metabolomic data in the context of human metabolic networks. We have developed a metabolite database by extracting and integrating information from several public sources. By querying this database, Metscape allows users to trace the connections between metabolites and genes, visualize compound networks and display compound structures as well as information for reactions, enzymes, genes and pathways. Applying the pathway filter, users can create subnetworks that consist of compounds and reactions from a given pathway. Metscape allows users to upload experimental data, and visualize and explore compound networks over time, or experimental conditions. Color and size of the nodes are used to visualize these dynamic changes. Metscape can display the entire metabolic network or any of the pathway-specific networks that exist in the database. AVAILABILITY: Metscape can be installed from within Cytoscape 2.6.x under 'Network and Attribute I/O' category. For more information, please visit http://metscape.ncibi.org/tryplugin.html.
V. Glenn Tarcea, Alla Karnovsky, Barbara Mirel, Terry E. Weymouth, Christopher W. Beecher, James D. Cavalcoli, Brian D. Athey, Gilbert S. Omenn, Charles F. Burant, H. V. Jagadish
Bioinform.9
2010 ConceptGen: a gene set enrichment and gene set relation mapping tool
abstract
Abstract Motivation: The elucidation of biological concepts enriched with differentially expressed genes has become an integral part of the analysis and interpretation of genomic data. Of additional importance is the ability to explore networks of relationships among previously defined biological concepts from diverse information sources, and to explore results visually from multiple perspectives. Accomplishing these tasks requires a unified framework for agglomeration of data from various genomic resources, novel visualizations, and user functionality. Results: We have developed ConceptGen, a web-based gene set enrichment and gene set relation mapping tool that is streamlined and simple to use. ConceptGen offers over 20 000 concepts comprising 14 different types of biological knowledge, including data not currently available in any other gene set enrichment or gene set relation mapping tool. We demonstrate the functionalities of ConceptGen using gene expression data modeling TGF-beta-induced epithelial-mesenchymal transition and metabolomics data comparing metastatic versus localized prostate cancers. Availability: ConceptGen is part of the NIH's National Center for Integrative Biomedical Informatics (NCIBI) and is freely available at http://conceptgen.ncibi.org. For terms of use, visit http://portal.ncibi.org/gateway/pdf/Terms%20of%20use-web.pdf Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Maureen A. Sartor, Vasudeva Mahavisno, Venkateshwar G. Keshamouni, James D. Cavalcoli, Zachary C. Wright, Alla Karnovsky, Rork Kuick, H. V. Jagadish, Barbara Mirel, Terry E. Weymouth, Brian D. Athey, Gilbert S. Omenn
Bioinform.12
2010 WaveletQuant, an improved quantification software based on wavelet signal threshold de-noising for labeled quantitative proteomic analysis
abstract
BACKGROUND: Quantitative proteomics technologies have been developed to comprehensively identify and quantify proteins in two or more complex samples. Quantitative proteomics based on differential stable isotope labeling is one of the proteomics quantification technologies. Mass spectrometric data generated for peptide quantification are often noisy, and peak detection and definition require various smoothing filters to remove noise in order to achieve accurate peptide quantification. Many traditional smoothing filters, such as the moving average filter, Savitzky-Golay filter and Gaussian filter, have been used to reduce noise in MS peaks. However, limitations of these filtering approaches often result in inaccurate peptide quantification. Here we present the WaveletQuant program, based on wavelet theory, for better or alternative MS-based proteomic quantification. RESULTS: We developed a novel discrete wavelet transform (DWT) and a 'Spatial Adaptive Algorithm' to remove noise and to identify true peaks. We programmed and compiled WaveletQuant using Visual C++ 2005 Express Edition. We then incorporated the WaveletQuant program in the Trans-Proteomic Pipeline (TPP), a commonly used open source proteomics analysis pipeline. CONCLUSIONS: We showed that WaveletQuant was able to quantify more proteins and to quantify them more accurately than the ASAPRatio, a program that performs quantification in the TPP pipeline, first using known mixed ratios of yeast extracts and then using a data set from ovarian cancer cell lysates. The program and its documentation can be downloaded from our website at http://systemsbiozju.org/data/WaveletQuant.
Qun Mo, David R. Goodlett, Leroy Hood, Gilbert S. Omenn, Biaoyang Lin
BMC Bioinform.6
2008 A compatible exon-exon junction database for the identification of exon skipping events using tandem mass spectrum data
abstract
BACKGROUND: Alternative splicing is an important gene regulation mechanism. It is estimated that about 74% of multi-exon human genes have alternative splicing. High throughput tandem (MS/MS) mass spectrometry provides valuable information for rapidly identifying potentially novel alternatively-spliced protein products from experimental datasets. However, the ability to identify alternative splicing events through tandem mass spectrometry depends on the database against which the spectra are searched. RESULTS: We wrote scripts in perl, Bioperl, mysql and Ensembl API and built a theoretical exon-exon junction protein database to account for all possible combinations of exons for a gene while keeping the frame of translation (i.e., keeping only in-phase exon-exon combinations) from the Ensembl Core Database. Using our liver cancer MS/MS dataset, we identified a total of 488 non-redundant peptides that represent putative exon skipping events. CONCLUSION: Our exon-exon junction database provides the scientific community with an efficient means to identify novel alternatively spliced (exon skipping) protein isoforms using mass spectrometry data. This database will be useful in annotating genome structures using rapidly accumulating proteomics data.
Xu Hong, Gilbert S. Omenn, Biaoyang Lin
BMC Bioinform.6