Sean D. Mooney

dblp:03/1985 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
5since 2021 · last 2023
0000-0003-2654-0833ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2023 Evaluation of crowdsourced mortality prediction models as a framework for assessing artificial intelligence in medicine
abstract
OBJECTIVE: Applications of machine learning in healthcare are of high interest and have the potential to improve patient care. Yet, the real-world accuracy of these models in clinical practice and on different patient subpopulations remains unclear. To address these important questions, we hosted a community challenge to evaluate methods that predict healthcare outcomes. We focused on the prediction of all-cause mortality as the community challenge question. MATERIALS AND METHODS: Using a Model-to-Data framework, 345 registered participants, coalescing into 25 independent teams, spread over 3 continents and 10 countries, generated 25 accurate models all trained on a dataset of over 1.1 million patients and evaluated on patients prospectively collected over a 1-year observation of a large health system. RESULTS: The top performing team achieved a final area under the receiver operator curve of 0.947 (95% CI, 0.942-0.951) and an area under the precision-recall curve of 0.487 (95% CI, 0.458-0.499) on a prospectively collected patient cohort. DISCUSSION: Post hoc analysis after the challenge revealed that models differ in accuracy on subpopulations, delineated by race or gender, even when they are trained on the same data. CONCLUSION: This is the largest community challenge focused on the evaluation of state-of-the-art machine learning methods in a healthcare system performed to date, revealing both opportunities and pitfalls of clinical AI.
Timothy Bergquist, Thomas Schaffter, Thomas Yu, Justin Prosser, Jifan Gao, Guanhua Chen 0002, Lukasz Charzewski, Zofia Nawalany, Ivan Brugere, Renata Retkute, Alidivinas Prusokas, Augustinas Prusokas, Yonghwa Choi, Junseok Choe, Inggeol Lee, Sunkyu Kim, Jaewoo Kang, Sean D. Mooney, Justin Guinney
J. Am. Medical Informatics Assoc.20
2022 Assessing Machine Learning Based Generators for Synthetic Electronic Health Records: A Benchmarking
Chao Yan 0004, Ziqi Zhang 0005, Zhiyu Wan, Justin Guinney, Sean D. Mooney, Bradley A. Malin
AMIA6
2022 Geospatial divide in real-world EHR data: Analytical workflow to assess regional biases and potential impact on health equity
Serena Jinchen Xie, Flavia Kapos, Stephen J. Mooney, Sean D. Mooney, Kari A. Stephens, Andrea L. Hartzler, Abhishek Pratap
AMIA4
2021 A Continuous Crowd-sourced Challenge for Benchmarking COVID-19 Health Outcome Prediction
Thomas Schaffter, Timothy Bergquist, Thomas Yu, Justin Prosser, Sean D. Mooney, Justin Guinney
AMIA6
2021 Information needs and priority use cases of population health researchers to improve preparedness for future hurricanes and floods
abstract
OBJECTIVE: Information gaps that accompany hurricanes and floods limit researchers' ability to determine the impact of disasters on population health. Defining key use cases for sharing complex disaster data with research communities and facilitators, and barriers to doing so are key to promoting population health research for disaster recovery. MATERIALS AND METHODS: We conducted a mixed-methods needs assessment with 15 population health researchers using interviews and card sorting. Interviews examined researchers' information needs by soliciting barriers and facilitators in the context of their expertise and research practices. Card sorting ranked priority use cases for disaster preparedness. RESULTS: Seven barriers and 6 facilitators emerged from interviews. Barriers to collaborative research included process limitations, collaboration dynamics, and perception of research importance. Barriers to data and technology adoption included data gaps, limitations in information quality, transparency issues, and difficulty to learn. Facilitators to collaborative research included collaborative engagement and human resource processes. Facilitators to data and technology adoption included situation awareness, data quality considerations, adopting community standards, and attractive to learn. Card sorting prioritized 15 use cases and identified 30 additional information needs for population health research in disaster preparedness. CONCLUSIONS: Population health researchers experience barriers to collaboration and adoption of data and technology that contribute to information gaps and limit disaster preparedness. The priority use cases we identified can help address information gaps by informing the design of supportive research tools and practices for disaster preparedness. Supportive tools should include information on data collection practices, quality assurance, and education resources usable during failures in electric or telecommunications systems.
Jimmy Phuong, Christina Bandaragoda, Shefali Haldar, Kari A. Stephens, Patricia Ordóñez 0002, Sean D. Mooney, Andrea L. Hartzler
J. Am. Medical Informatics Assoc.6
2020 Piloting a model-to-data approach to enable predictive analytics in health care through patient mortality prediction
abstract
OBJECTIVE: The development of predictive models for clinical application requires the availability of electronic health record (EHR) data, which is complicated by patient privacy concerns. We showcase the "Model to Data" (MTD) approach as a new mechanism to make private clinical data available for the development of predictive models. Under this framework, we eliminate researchers' direct interaction with patient data by delivering containerized models to the EHR data. MATERIALS AND METHODS: We operationalize the MTD framework using the Synapse collaboration platform and an on-premises secure computing environment at the University of Washington hosting EHR data. Containerized mortality prediction models developed by a model developer, were delivered to the University of Washington via Synapse, where the models were trained and evaluated. Model performance metrics were returned to the model developer. RESULTS: The model developer was able to develop 3 mortality prediction models under the MTD framework using simple demographic features (area under the receiver-operating characteristic curve [AUROC], 0.693), demographics and 5 common chronic diseases (AUROC, 0.861), and the 1000 most common features from the EHR's condition/procedure/drug domains (AUROC, 0.921). DISCUSSION: We demonstrate the feasibility of the MTD framework to facilitate the development of predictive models on private EHR data, enabled by common data models and containerization software. We identify challenges that both the model developer and the health system information technology group encountered and propose future efforts to improve implementation. CONCLUSIONS: The MTD framework lowers the barrier of access to EHR data and can accelerate the development and evaluation of clinical prediction models.
Timothy Bergquist, Thomas Schaffter, Thomas Yu, Vikas Pejaver, Noah Hammarlund, Justin Prosser, Justin Guinney, Sean D. Mooney
J. Am. Medical Informatics Assoc.9
2020 Leaf: an open-source, model-agnostic, data-driven web application for cohort discovery and translational biomedical research
abstract
OBJECTIVE: Academic medical centers and health systems are increasingly challenged with supporting appropriate secondary use of clinical data. Enterprise data warehouses have emerged as central resources for these data, but often require an informatician to extract meaningful information, limiting direct access by end users. To overcome this challenge, we have developed Leaf, a lightweight self-service web application for querying clinical data from heterogeneous data models and sources. MATERIALS AND METHODS: Leaf utilizes a flexible biomedical concept system to define hierarchical concepts and ontologies. Each Leaf concept contains both textual representations and SQL query building blocks, exposed by a simple drag-and-drop user interface. Leaf generates abstract syntax trees which are compiled into dynamic SQL queries. RESULTS: Leaf is a successful production-supported tool at the University of Washington, which hosts a central Leaf instance querying an enterprise data warehouse with over 300 active users. Through the support of UW Medicine (https://uwmedicine.org), the Institute of Translational Health Sciences (https://www.iths.org), and the National Center for Data to Health (https://ctsa.ncats.nih.gov/cd2h/), Leaf source code has been released into the public domain at https://github.com/uwrit/leaf. DISCUSSION: Leaf allows the querying of single or multiple clinical databases simultaneously, even those of different data models. This enables fast installation without costly extraction or duplication. CONCLUSIONS: Leaf differs from existing cohort discovery tools because it does not specify a required data model and is designed to seamlessly leverage existing user authentication systems and clinical databases in situ. We believe Leaf to be useful for health system analytics, clinical research data warehouses, precision medicine biobanks, and clinical studies involving large patient cohorts.
Nicholas J. Dobbins, Clifford H. Spital, Robert A. Black, Jason M. Morrison, Bas de Veer, Elizabeth Zampino, Robert D. Harrington, Bethene D. Britt, Kari A. Stephens, Adam B. Wilcox, Peter Tarczy-Hornoch, Sean D. Mooney
J. Am. Medical Informatics Assoc.12
2019 Detecting Seasonal, Holiday, and Rare Events from Trauma Data in the Electronic Health Record
Timothy Bergquist, Vikas Pejaver, Noah Hammarlund, Sean D. Mooney, Stephen J. Mooney
AMIA4
2019 Continuing challenges swirl around bioinformatics service delivery
Sean D. Mooney
J. Biomed. Informatics1
2019 Pathogenicity and functional impact of non-frameshifting insertion/deletion variation in the human genome
abstract
Differentiation between phenotypically neutral and disease-causing genetic variation remains an open and relevant problem. Among different types of variation, non-frameshifting insertions and deletions (indels) represent an understudied group with widespread phenotypic consequences. To address this challenge, we present a machine learning method, MutPred-Indel, that predicts pathogenicity and identifies types of functional residues impacted by non-frameshifting insertion/deletion variation. The model shows good predictive performance as well as the ability to identify impacted structural and functional residues including secondary structure, intrinsic disorder, metal and macromolecular binding, post-translational modifications, allosteric sites, and catalytic residues. We identify structural and functional mechanisms impacted preferentially by germline variation from the Human Gene Mutation Database, recurrent somatic variation from COSMIC in the context of different cancers, as well as de novo variants from families with autism spectrum disorder. Further, the distributions of pathogenicity prediction scores generated by MutPred-Indel are shown to differentiate highly recurrent from non-recurrent somatic variation. Collectively, we present a framework to facilitate the interrogation of both pathogenicity and the functional effects of non-frameshifting insertion/deletion variants. The MutPred-Indel webserver is available at http://mutpred.mutdb.org/.
Kymberleigh A. Pagel, Danny Antaki, Aojie Lian, Matthew E. Mort, David N. Cooper, Jonathan Sebat, Lilia M. Iakoucheva, Sean D. Mooney, Predrag Radivojac
PLoS Comput. Biol.8
2018 Uncovering exposures responsible for birth season - disease effects: a global study
abstract
OBJECTIVE: Birth month and climate impact lifetime disease risk, while the underlying exposures remain largely elusive. We seek to uncover distal risk factors underlying these relationships by probing the relationship between global exposure variance and disease risk variance by birth season. MATERIAL AND METHODS: This study utilizes electronic health record data from 6 sites representing 10.5 million individuals in 3 countries (United States, South Korea, and Taiwan). We obtained birth month-disease risk curves from each site in a case-control manner. Next, we correlated each birth month-disease risk curve with each exposure. A meta-analysis was then performed of correlations across sites. This allowed us to identify the most significant birth month-exposure relationships supported by all 6 sites while adjusting for multiplicity. We also successfully distinguish relative age effects (a cultural effect) from environmental exposures. RESULTS: Attention deficit hyperactivity disorder was the only identified relative age association. Our methods identified several culprit exposures that correspond well with the literature in the field. These include a link between first-trimester exposure to carbon monoxide and increased risk of depressive disorder (R = 0.725, confidence interval [95% CI], 0.529-0.847), first-trimester exposure to fine air particulates and increased risk of atrial fibrillation (R = 0.564, 95% CI, 0.363-0.715), and decreased exposure to sunlight during the third trimester and increased risk of type 2 diabetes mellitus (R = -0.816, 95% CI, -0.5767, -0.929). CONCLUSION: A global study of birth month-disease relationships reveals distal risk factors involved in causal biological pathways that underlie them.
Mary Regina Boland, Pradipta Parhi, Li Li 0062, Riccardo Miotto, Robert J. Carroll, Usman Iqbal, Phung Anh Nguyen, Martijn J. Schuemie, Seng Chan You, Donahue Smith, Sean D. Mooney, Patrick B. Ryan, Yu-Chuan Li, Rae Woong Park, Joshua C. Denny, Joel Dudley, George Hripcsak, Pierre Gentine, Nicholas P. Tatonetti
J. Am. Medical Informatics Assoc.11
2017 When loss-of-function is loss of function: assessing mutational signatures and impact of loss-of-function genetic variants
abstract
MOTIVATION: Loss-of-function genetic variants are frequently associated with severe clinical phenotypes, yet many are present in the genomes of healthy individuals. The available methods to assess the impact of these variants rely primarily upon evolutionary conservation with little to no consideration of the structural and functional implications for the protein. They further do not provide information to the user regarding specific molecular alterations potentially causative of disease. RESULTS: To address this, we investigate protein features underlying loss-of-function genetic variation and develop a machine learning method, MutPred-LOF, for the discrimination of pathogenic and tolerated variants that can also generate hypotheses on specific molecular events disrupted by the variant. We investigate a large set of human variants derived from the Human Gene Mutation Database, ClinVar and the Exome Aggregation Consortium. Our prediction method shows an area under the Receiver Operating Characteristic curve of 0.85 for all loss-of-function variants and 0.75 for proteins in which both pathogenic and neutral variants have been observed. We applied MutPred-LOF to a set of 1142 de novo vari3ants from neurodevelopmental disorders and find enrichment of pathogenic variants in affected individuals. Overall, our results highlight the potential of computational tools to elucidate causal mechanisms underlying loss of protein function in loss-of-function variants. AVAILABILITY AND IMPLEMENTATION: http://mutpred.mutdb.org. CONTACT: [email protected].
Kymberleigh A. Pagel, Vikas Pejaver, Guan Ning Lin, Hyun-Jun Nam, Matthew E. Mort, David N. Cooper, Jonathan Sebat, Lilia M. Iakoucheva, Sean D. Mooney, Predrag Radivojac
Bioinform.9
2016 The Loss and Gain of Functional Amino Acid Residues Is a Common Mechanism Causing Human Inherited Disease
abstract
Elucidating the precise molecular events altered by disease-causing genetic variants represents a major challenge in translational bioinformatics. To this end, many studies have investigated the structural and functional impact of amino acid substitutions. Most of these studies were however limited in scope to either individual molecular functions or were concerned with functional effects (e.g. deleterious vs. neutral) without specifically considering possible molecular alterations. The recent growth of structural, molecular and genetic data presents an opportunity for more comprehensive studies to consider the structural environment of a residue of interest, to hypothesize specific molecular effects of sequence variants and to statistically associate these effects with genetic disease. In this study, we analyzed data sets of disease-causing and putatively neutral human variants mapped to protein 3D structures as part of a systematic study of the loss and gain of various types of functional attribute potentially underlying pathogenic molecular alterations. We first propose a formal model to assess probabilistically function-impacting variants. We then develop an array of structure-based functional residue predictors, evaluate their performance, and use them to quantify the impact of disease-causing amino acid substitutions on catalytic activity, metal binding, macromolecular binding, ligand binding, allosteric regulation and post-translational modifications. We show that our methodology generates actionable biological hypotheses for up to 41% of disease-causing genetic variants mapped to protein structures suggesting that it can be reliably used to guide experimental validation. Our results suggest that a significant fraction of disease-causing human variants mapping to protein structures are function-altering both in the presence and absence of stability disruption.
Jose Lugo-Martinez, Vikas Pejaver, Kymberleigh A. Pagel, Shantanu Jain, Matthew E. Mort, David N. Cooper, Sean D. Mooney, Predrag Radivojac
PLoS Comput. Biol.7
2015 Ten Simple Rules for a Community Computational Challenge
abstract
In science, the relationship between methods and discovery is symbiotic.As we discover more, we are able to construct more precise and sensitive tools and methods that enable further discovery.With better lens crafting came microscopes, and with them the discovery of living cells.In the last 40 years, advances in molecular biology, statistics, and computer science have ushered in the field of bioinformatics and the genomic era.Computational scientists enjoy developing new methods, and the community encourages them to do so.Indeed, the editorial guidelines for PLOS Computational Biology require manuscripts to apply novel methods.However, it is often confusing to know which method to choose: which method is best?And, in this context, what does "best" mean?To help choose an appropriate method for a particular task, scientists often form community-based challenges for the unbiased evaluation of methods in a given field.These challenges help evaluate existing and novel methods, while helping to coalesce a community and leading to new ideas and collaborations.In computational biology, the first of these challenges was arguably the Critical Assessment of protein Structure Prediction, or CASP [1], whose goal is to evaluate methods for predicting three-dimensional protein structure from amino acid sequence.The first CASP meeting was held in December of 1994, following a "prediction period" where members of the community were presented with protein amino acid sequences and asked to predict their three dimensional structures.The sequences that were chosen had recently been solved by X-ray crystallography but had not been not published or released until after the predictions from the community were made.Since the first CASP, we have seen many successful challenges, including Critical Assessment of Function Annotation (CAFA) for protein function prediction [2], Critical Assessment of Genome Interpretation (CAGI) (for genome interpretation) [3], Critical Assessment of Massive (originally "Microarray") Data Analysis (CAMDA) (for large-scale biological data) [4], BioCreative (for biomedical text mining) [5], the Assemblathon (for sequence assembly), and the NCI-DREAM Challenges (for various biomedical challenges), amongst others [6].Computational challenges also help solve new problems.While the original CASP experiment was developed to evaluate existing methods applied to current problems, other communities often look at other areas for which there are no existing tools.These challenges have spread successfully to industry, and companies such as Innocentive [7] and X-Prize [8] offer large prizes for solving novel questions.
Iddo Friedberg, Mark N. Wass, Sean D. Mooney, Predrag Radivojac
PLoS Comput. Biol.3
2014 The automated function prediction SIG looks back at 2013 and prepares for 2014
abstract
Mark N. Wass, Sean D. Mooney, Michal Linial, Predrag Radivojac, Iddo Friedberg
Bioinform.2
2014 A Probabilistic Model to Predict Clinical Phenotypic Traits from Genome Sequencing
abstract
Genetic screening is becoming possible on an unprecedented scale. However, its utility remains controversial. Although most variant genotypes cannot be easily interpreted, many individuals nevertheless attempt to interpret their genetic information. Initiatives such as the Personal Genome Project (PGP) and Illumina's Understand Your Genome are sequencing thousands of adults, collecting phenotypic information and developing computational pipelines to identify the most important variant genotypes harbored by each individual. These pipelines consider database and allele frequency annotations and bioinformatics classifications. We propose that the next step will be to integrate these different sources of information to estimate the probability that a given individual has specific phenotypes of clinical interest. To this end, we have designed a Bayesian probabilistic model to predict the probability of dichotomous phenotypes. When applied to a cohort from PGP, predictions of Gilbert syndrome, Graves' disease, non-Hodgkin lymphoma, and various blood groups were accurate, as individuals manifesting the phenotype in question exhibited the highest, or among the highest, predicted probabilities. Thirty-eight PGP phenotypes (26%) were predicted with area-under-the-ROC curve (AUC)>0.7, and 23 (15.8%) of these were statistically significant, based on permutation tests. Moreover, in a Critical Assessment of Genome Interpretation (CAGI) blinded prediction experiment, the models were used to match 77 PGP genomes to phenotypic profiles, generating the most accurate prediction of 16 submissions, according to an independent assessor. Although the models are currently insufficiently accurate for diagnostic utility, we expect their performance to improve with growth of publicly available genomics data and model refinement by domain experts.
Yun-Ching Chen, Christopher Douville, Noushin Niknafs, Grace H. T. Yeo, Violeta Beleva Guthrie, Hannah Carter, Peter D. Stenson, David N. Cooper, Sean D. Mooney, Rachel Karchin
PLoS Comput. Biol.11
2013 STOP using just GO: a multi-ontology hypothesis generation tool for high throughput experimentation
abstract
BACKGROUND: Gene Ontology (GO) enrichment analysis remains one of the most common methods for hypothesis generation from high throughput datasets. However, we believe that researchers strive to test other hypotheses that fall outside of GO. Here, we developed and evaluated a tool for hypothesis generation from gene or protein lists using ontological concepts present in manually curated text that describes those genes and proteins. RESULTS: As a consequence we have developed the method Statistical Tracking of Ontological Phrases (STOP) that expands the realm of testable hypotheses in gene set enrichment analyses by integrating automated annotations of genes to terms from over 200 biomedical ontologies. While not as precise as manually curated terms, we find that the additional enriched concepts have value when coupled with traditional enrichment analyses using curated terms. CONCLUSION: Multiple ontologies have been developed for gene and protein annotation, by using a dataset of both manually curated GO terms and automatically recognized concepts from curated text we can expand the realm of hypotheses that can be discovered. The web application STOP is available at http://mooneygroup.org/stop/.
Tobias Wittkop, Emily TerAvest, Uday S. Evani, K. Mathew Fleisch, Ari E. Berman, Corey Powell, Nigam H. Shah, Sean D. Mooney
BMC Bioinform.8
2011 Identifying viral integration sites using SeqMap 2.0
abstract
UNLABELLED: Retroviral integration has been implicated in several biomedical applications, including identification of cancer-associated genes and malignant transformation in gene therapy clinical trials. We introduce an efficient and scalable method for fast identification of viral vector integration sites from long read high-throughput sequencing. Individual sequence reads are masked to remove non-genomic sequence, aligned to the host genome and assembled into contiguous fragments used to pinpoint the position of integration. AVAILABILITY AND IMPLEMENTATION: The method is implemented in a publicly accessible web server platform, SeqMap 2.0, containing analysis tools and both private and shared lab workspaces that facilitate collaboration among researchers. Available at http://seqmap.compbio.iupui.edu/.
Troy B. Hawkins, Jessica Dantzer, Brandon Peters, Mary Dinauer, Keithanne Mockaitis, Sean D. Mooney, Kenneth Cornetta
Bioinform.6
2010 Structure-based kernels for the prediction of catalytic residues and their involvement in human inherited disease
abstract
MOTIVATION: Enzyme catalysis is involved in numerous biological processes and the disruption of enzymatic activity has been implicated in human disease. Despite this, various aspects of catalytic reactions are not completely understood, such as the mechanics of reaction chemistry and the geometry of catalytic residues within active sites. As a result, the computational prediction of catalytic residues has the potential to identify novel catalytic pockets, aid in the design of more efficient enzymes and also predict the molecular basis of disease. RESULTS: We propose a new kernel-based algorithm for the prediction of catalytic residues based on protein sequence, structure and evolutionary information. The method relies upon explicit modeling of similarity between residue-centered neighborhoods in protein structures. We present evidence that this algorithm evaluates favorably against established approaches, and also provides insights into the relative importance of the geometry, physicochemical properties and evolutionary conservation of catalytic residue activity. The new algorithm was used to identify known mutations associated with inherited disease whose molecular mechanism might be predicted to operate specifically though the loss or gain of catalytic residues. It should, therefore, provide a viable approach to identifying the molecular basis of disease in which the loss or gain of function is not caused solely by the disruption of protein stability. Our analysis suggests that both mechanisms are actively involved in human inherited disease. AVAILABILITY AND IMPLEMENTATION: Source code for the structural kernel is available at www.informatics.indiana.edu/predrag/.
Fuxiao Xin, Steven Myers, Yong Fuga Li, David N. Cooper, Sean D. Mooney, Predrag Radivojac
Bioinform.5
2010 Structure-based kernels for the prediction of catalytic residues and their involvement in human inherited disease
abstract
Enzyme catalysis is involved in numerous biological processes and the disruption of enzymatic activity has been implicated in human disease. Despite the functional importance, various aspects of catalytic reactions are not completely understood, such as the mechanics of reaction chemistry and the geometry of catalytic residues within active sites. As a result, the computational prediction of catalytic residues has the potential to identify novel catalytic pockets, aid in the design of more efficient enzymes and also predict the molecular basis of disease. We proposed a new kernel-based algorithm for the prediction of catalytic residues and functional sites in general in protein structures [ 1 ]. The method relies upon explicit modelling of similarity between residue-centred neighbourhoods in protein structures. Specifically, we start with a construction of oriented structural neighbourhoods followed by separating the neighbourhood volume into small cells. The similarities between two structural neighbourhoods are accumulation of their similarity in each cell. The kernel function is a product of three kernels, each addressing a separate aspect of protein function: (i) the geometric kernel addresses the shape similarity, (ii) the chemical kernel addresses the similarity in physicochemical properties, and (iii) the evolutionary kernel addresses the evolutionary similarity of conservation patterns for the residues in two structural neighbourhoods. Our approach was favourably evaluated against two of the leading alternative approaches, FEATURE [ 2 ] and GBT [ 3 ], as shown in Table 1 . The new algorithm was used to identify known mutations associated with inherited disease whose molecular mechanism might be predicted to operate specifically though the loss or gain of catalytic residues. It should therefore provide a viable approach in identifying the molecular basis of disease in which the loss or gain of function is not caused solely by the disruption of protein stability. Our analysis suggests that both loss and gain of catalytic residues are actively involved in human inherited disease. Our kernel method for functional sites prediction based on protein structures evaluates favourably against established methods on the same data set using the same evaluation procedure. The results from applying our catalytic residue predictor to disease mutations indicated that both loss and gain of catalytic residues are actively involved in human inherited disease.
Fuxiao Xin, Steven Myers, Yong Fuga Li, David N. Cooper, Sean D. Mooney, Predrag Radivojac
BMC Bioinform.5
2009 Automated inference of molecular mechanisms of disease from amino acid substitutions
abstract
MOTIVATION: Advances in high-throughput genotyping and next generation sequencing have generated a vast amount of human genetic variation data. Single nucleotide substitutions within protein coding regions are of particular importance owing to their potential to give rise to amino acid substitutions that affect protein structure and function which may ultimately lead to a disease state. Over the last decade, a number of computational methods have been developed to predict whether such amino acid substitutions result in an altered phenotype. Although these methods are useful in practice, and accurate for their intended purpose, they are not well suited for providing probabilistic estimates of the underlying disease mechanism. RESULTS: We have developed a new computational model, MutPred, that is based upon protein sequence, and which models changes of structural features and functional sites between wild-type and mutant sequences. These changes, expressed as probabilities of gain or loss of structure and function, can provide insight into the specific molecular mechanism responsible for the disease state. MutPred also builds on the established SIFT method but offers improved classification accuracy with respect to human disease mutations. Given conservative thresholds on the predicted disruption of molecular function, we propose that MutPred can generate accurate and reliable hypotheses on the molecular basis of disease for approximately 11% of known inherited disease-causing mutations. We also note that the proportion of changes of functionally relevant residues in the sets of cancer-associated somatic mutations is higher than for the inherited lesions in the Human Gene Mutation Database which are instead predicted to be characterized by disruptions of protein structure. AVAILABILITY: http://mutdb.org/mutpred CONTACT: [email protected]; [email protected].
Vidhya G. Krishnan, Matthew E. Mort, Fuxiao Xin, Kishore K. Kamati, David N. Cooper, Sean D. Mooney, Predrag Radivojac
Bioinform.7
2008 Extensible open source content management systems and frameworks: a solution for many needs of a bioinformatics group
abstract
A common challenge for bioinformaticians, in either academic or industry laboratory environments, is providing informatic solutions via the Internet or through a web browser. Recently, the open source community began developing tools for building and maintaining web applications for many disciplines. These content management systems (CMS) provide many of the basic needs of an informatics group, whether in a small company, a group within a larger organisation or an academic laboratory. These tools aid in managing software development, website development, document development, course development, datasets, collaborations and customers. Since many of these tools are extensible, they can be developed to support other research-specific activities, such as handling large biomedical datasets or deploying bioanalytic tools. In this review of open source website management tools, the basic features of content management systems are discussed along with commonly used open source software. Additionally, some examples of their use in biomedical research are given.
Sean D. Mooney, Peter H. Baenziger
Briefings Bioinform.1
2005 Bioinformatics approaches and resources for single nucleotide polymorphism functional analysis
abstract
Since the initial sequencing of the human genome, many projects are underway to understand the effects of genetic variation between individuals. Predicting and understanding the downstream effects of genetic variation using computational methods are becoming increasingly important for single nucleotide polymorphism (SNP) selection in genetics studies and understanding the molecular basis of disease. According to the NIH, there are now more than four million validated SNPs in the human genome. The volume of known genetic variations lends itself well to an informatics approach. Bioinformaticians have become very good at functional inference methods derived from functional and structural genomics. This review will present a broad overview of the tools and resources available to collect and understand functional variation from the perspective of structure, expression, evolution and phenotype. Additionally, public resources available for SNP identification and characterisation are summarised.
Sean D. Mooney
Briefings Bioinform.1
2003 MutDB: annotating human variation with functionally relevant data
abstract
SUMMARY: We have developed a resource, MutDB (http://mutdb.org/), to aid in determining which single nucleotide polymorphisms (SNPs) are likely to alter the function of their associated protein product. MutDB contains protein structure annotations and comparative genomic annotations for 8000 disease-associated mutations and SNPs found in the UCSC Annotated Genome and the human RefSeq gene set. MutDB provides interactive mutation maps at the gene and protein levels, and allows for ranking of their predicted functional consequences based on conservation in multiple sequence alignments. AVAILABILITY: http://mutdb.org/ SUPPLEMENTARY INFORMATION: http://mutdb.org/about/about.html
Sean D. Mooney, Russ B. Altman
Bioinform.1
2003 Analysis of Mutations in the COLIA1 Gene with Second-Order Rule Induction
abstract
Mutations are structural changes in DNA that can cause protein malfunction and genetic disease. This paper describes a machine learning approach to analyzing mutations associated to Osteogenesis Imperfecta (OI), also known as brittle bone disease. We apply SORCER, a second-order rule induction system to predict clinical phenotypes of OI from mutation and neighboring amino acid sequences in the COLIA1 gene. On the average, over ten 10-fold cross-validations, SORCER gives more accurate results than C4.5 with average accuracy of about 81.2%. The paper discusses the advantages and limitations of SORCER and demonstrates its use to provide initial exploration of biological sequences.
Rattikorn Hewett, John H. Leuchner, Sean D. Mooney, Teri E. Klein
Int. J. Pattern Recognit. Artif. Intell.3
2002 The functional importance of disease-associated mutation
abstract
BACKGROUND: For many years, scientists believed that point mutations in genes are the genetic switches for somatic and inherited diseases such as cystic fibrosis, phenylketonuria and cancer. Some of these mutations likely alter a protein's function in a manner that is deleterious, and they should occur in functionally important regions of the protein products of genes. Here we show that disease-associated mutations occur in regions of genes that are conserved, and can identify likely disease-causing mutations. RESULTS: To show this, we have determined conservation patterns for 6185 non-synonymous and heritable disease-associated mutations in 231 genes. We define a parameter, the conservation ratio, as the ratio of average negative entropy of analyzable positions with reported mutations to that of every analyzable position in the gene sequence. We found that 84.0% of the 231 genes have conservation ratios less than one. 139 genes had eleven or more analyzable mutations and 88.0% of those had conservation ratios less than one. CONCLUSIONS: These results indicate that phylogenetic information is a powerful tool for the study of disease-associated mutations. Our alignments and analysis has been made available as part of the database at http://cancer.stanford.edu/mut-paper/. Within this dataset, each position is annotated with the analysis, so the most likely disease-causing mutations can be identified.
Sean D. Mooney, Teri E. Klein
BMC Bioinform.1