EDBT 2026 Demo / reviewers in the wild / expert
Tom R. Gaunt
dblp:63/2162
· DBLP profile ↗
27ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-0924-3247ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Integrating Mendelian randomization and literature-mined evidence for breast cancer risk factorsabstractOBJECTIVE: An increasing challenge in population health research is efficiently utilising the wealth of data available from multiple sources to investigate disease mechanisms and identify potential intervention targets. The use of biomedical data integration platforms can facilitate evidence triangulation from these different sources, improving confidence in causal relationships of interest. In this work, we aimed to integrate Mendelian randomization (MR) and literature-mined evidence from the EpiGraphDB biomedical knowledge graph to build a comprehensive overview of risk factors for developing breast cancer. METHODS: We utilised MR-EvE ("Everything-vs-Everything") data to identify candidate risk factors for breast cancer and generate hypotheses for potential mediators of their effect. We also integrated this data with literature-mined relationships, which were extracted by overlapping literature spaces of risk factors and breast cancer. The literature-based discovery (LBD) results were followed up by validation with two-step MR to triangulate the findings from two data sources. RESULTS: We identified 129 novel and established lifestyle risk factors and molecular traits with evidence of an effect on breast cancer, and made the MR results available in an R/Shiny app (https://mvab.shinyapps.io/MR_heatmaps/). We developed an LBD approach for identifying potential mechanistic intermediates of identified risk factors. We present the results of MR and literature evidence integration for two case studies (childhood body size and HDL-cholesterol), demonstrating their complementary functionalities. CONCLUSION: We demonstrate that MR-EvE data offers an efficient hypothesis-generating approach for identifying disease risk factors. Moreover, we show that integrating MR evidence with literature-mined data may be used to identify causal intermediates and uncover the mechanisms behind the disease. Marina Vabistsevits, Timothy Robinson, Benjamin L. Elsworth, Yi Liu 0058, Tom R. Gaunt |
J. Biomed. Informatics | 5 |
| 2024 | Triangulating evidence in health sciences with Annotated Semantic QueriesabstractMOTIVATION: Integrating information from data sources representing different study designs has the potential to strengthen evidence in population health research. However, this concept of evidence "triangulation" presents a number of challenges for systematically identifying and integrating relevant information. These include the harmonization of heterogenous evidence with common semantic concepts and properties, as well as the priortization of the retrieved evidence for triangulation with the question of interest. RESULTS: We present Annotated Semantic Queries (ASQ), a natural language query interface to the integrated biomedical entities and epidemiological evidence in EpiGraphDB, which enables users to extract "claims" from a piece of unstructured text, and then investigate the evidence that could either support, contradict the claims, or offer additional information to the query. This approach has the potential to support the rapid review of preprints, grant applications, conference abstracts, and articles submitted for peer review. ASQ implements strategies to harmonize biomedical entities in different taxonomies and evidence from different sources, to facilitate evidence triangulation and interpretation. AVAILABILITY AND IMPLEMENTATION: ASQ is openly available at https://asq.epigraphdb.org and its source code is available at https://github.com/mrcieu/epigraphdb-asq under GPL-3.0 license. Yi Liu 0058, Tom R. Gaunt |
Bioinform. | 2 |
| 2024 | DrivR-Base: a feature extraction toolkit for variant effect prediction model constructionabstractMOTIVATION: Recent advancements in sequencing technologies have led to the discovery of numerous variants in the human genome. However, understanding their precise roles in diseases remains challenging due to their complex functional mechanisms. Various methodologies have emerged to predict the pathogenic significance of these genetic variants. Typically, these methods employ an integrative approach, leveraging diverse data sources that provide important insights into genomic function. Despite the abundance of publicly available data sources and databases, the process of navigating, extracting, and pre-processing features for machine learning models can be highly challenging and time-consuming. Furthermore, researchers often invest substantial effort in feature extraction, only to later discover that these features lack informativeness. RESULTS: In this article, we introduce DrivR-Base, an innovative resource that efficiently extracts and integrates molecular information (features) related to single nucleotide variants. These features encompass information about the genomic positions and the associated protein positions of a variant. They are derived from a wide array of databases and tools, including structural properties obtained from AlphaFold, regulatory information sourced from ENCODE, and predicted variant consequences from Variant Effect Predictor. DrivR-Base is easily deployable via a Docker container to ensure reproducibility and ease of access across diverse computational environments. The resulting features can be used as input for machine learning models designed to predict the pathogenic impact of human genome variants in disease. Moreover, these feature sets have applications beyond this, including haploinsufficiency prediction and the development of drug repurposing tools. We describe the resource's development, practical applications, and potential for future expansion and enhancement. AVAILABILITY AND IMPLEMENTATION: DrivR-Base source code is available at https://github.com/amyfrancis97/DrivR-Base. Amy Francis, Colin Campbell, Tom R. Gaunt |
Bioinform. | 3 |
| 2024 | Fast polypharmacy side effect prediction using tensor factorizationabstractMOTIVATION: Adverse reactions from drug combinations are increasingly common, making their accurate prediction a crucial challenge in modern medicine. Laboratory-based identification of these reactions is insufficient due to the combinatorial nature of the problem. While many computational approaches have been proposed, tensor factorization (TF) models have shown mixed results, necessitating a thorough investigation of their capabilities when properly optimized. RESULTS: We demonstrate that TF models can achieve state-of-the-art performance on polypharmacy side effect prediction, with our best model (SimplE) achieving median scores of 0.978 area under receiver-operating characteristic curve, 0.971 area under precision-recall curve, and 1.000 AP@50 across 963 side effects. Notably, this model reaches 98.3% of its maximum performance after just two epochs of training (approximately 4 min), making it substantially faster than existing approaches while maintaining comparable accuracy. We also find that incorporating monopharmacy data as self-looping edges in the graph performs marginally better than using it to initialize embeddings. AVAILABILITY AND IMPLEMENTATION: All code used in the experiments is available in our GitHub repository (https://doi.org/10.5281/zenodo.10684402). The implementation was carried out using Python 3.8.12 with PyTorch 1.7.1, accelerated with CUDA 11.4 on NVIDIA GeForce RTX 2080 Ti GPUs. Oliver Lloyd, Yi Liu 0058, Tom R. Gaunt |
Bioinform. | 3 |
| 2023 | Using language models and ontology topology to perform semantic mapping of traits between biomedical datasetsabstractMOTIVATION: Human traits are typically represented in both the biomedical literature and large population studies as descriptive text strings. Whilst a number of ontologies exist, none of these perfectly represent the entire human phenome and exposome. Mapping trait names across large datasets is therefore time-consuming and challenging. Recent developments in language modelling have created new methods for semantic representation of words and phrases, and these methods offer new opportunities to map human trait names in the form of words and short phrases, both to ontologies and to each other. Here, we present a comparison between a range of established and more recent language modelling approaches for the task of mapping trait names from UK Biobank to the Experimental Factor Ontology (EFO), and also explore how they compare to each other in direct trait-to-trait mapping. RESULTS: In our analyses of 1191 traits from UK Biobank with manual EFO mappings, the BioSentVec model performed best at predicting these, matching 40.3% of the manual mappings correctly. The BlueBERT-EFO model (finetuned on EFO) performed nearly as well (38.8% of traits matching the manual mapping). In contrast, Levenshtein edit distance only mapped 22% of traits correctly. Pairwise mapping of traits to each other demonstrated that many of the models can accurately group similar traits based on their semantic similarity. AVAILABILITY AND IMPLEMENTATION: Our code is available at https://github.com/MRCIEU/vectology. Yi Liu 0058, Benjamin L. Elsworth, Tom R. Gaunt |
Bioinform. | 3 |
| 2021 | Prediction of driver variants in the cancer genome via machine learning methodologiesabstractSequencing technologies have led to the identification of many variants in the human genome which could act as disease-drivers. As a consequence, a variety of bioinformatics tools have been proposed for predicting which variants may drive disease, and which may be causatively neutral. After briefly reviewing generic tools, we focus on a subset of these methods specifically geared toward predicting which variants in the human cancer genome may act as enablers of unregulated cell proliferation. We consider the resultant view of the cancer genome indicated by these predictors and discuss ways in which these types of prediction tools may be progressed by further research. Mark F. Rogers, Tom R. Gaunt, Colin Campbell |
Briefings Bioinform. | 2 |
| 2021 | MELODI Presto: a fast and agile tool to explore semantic triples derived from biomedical literatureabstractSUMMARY: The field of literature-based discovery is growing in step with the volume of literature being produced. From modern natural language processing algorithms to high quality entity tagging, the methods and their impact are developing rapidly. One annotation object that arises from these approaches, the subject-predicate-object triple, is proving to be very useful in representing knowledge. We have implemented efficient search methods and an application programming interface, to create fast and convenient functions to utilize triples extracted from the biomedical literature by SemMedDB. By refining these data, we have identified a set of triples that focus on the mechanistic aspects of the literature, and provide simple methods to explore both enriched triples from single queries, and overlapping triples across two query lists. AVAILABILITY AND IMPLEMENTATION: https://melodi-presto.mrcieu.ac.uk/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin L. Elsworth, Tom R. Gaunt |
Bioinform. | 2 |
| 2021 | Erratum to: EpiGraphDB: a database and data mining platform for health data scienceabstractBioinformatics (2020) doi: 10.1093/bioinformatics/btaa961 Upon the original publication of this article, there was an error in the source code syntax under sub-section “2.2 Integration of epidemiological evidence” in the “Materials and methods” section. The source code syntax should read: “(e.g. (Gwas {trait: ‘Body mass index’})-[MR {beta, se, pval}]->(Gwas {trait: ‘Coronary heart disease’}))” instead of “ (e.g. [Gwas (trait: ‘Body mass index’)]-[MR {beta, se, pval}]->(Gwas {trait: ‘Coronary heart disease’})))”. This error has now been corrected. The Publisher apologizes for the error. Yi Liu 0058, Benjamin L. Elsworth, Pau Erola, Valeriia Haberland, Gibran Hemani, Matt Lyon, Jie Zheng 0006, Oliver Lloyd, Marina Vabistsevits, Tom R. Gaunt |
Bioinform. | 10 |
| 2021 | EpiGraphDB: a database and data mining platform for health data scienceabstractMOTIVATION: The wealth of data resources on human phenotypes, risk factors, molecular traits and therapeutic interventions presents new opportunities for population health sciences. These opportunities are paralleled by a growing need for data integration, curation and mining to increase research efficiency, reduce mis-inference and ensure reproducible research. RESULTS: We developed EpiGraphDB (https://epigraphdb.org/), a graph database containing an array of different biomedical and epidemiological relationships and an analytical platform to support their use in human population health data science. In addition, we present three case studies that illustrate the value of this platform. The first uses EpiGraphDB to evaluate potential pleiotropic relationships, addressing mis-inference in systematic causal analysis. In the second case study, we illustrate how protein-protein interaction data offer opportunities to identify new drug targets. The final case study integrates causal inference using Mendelian randomization with relationships mined from the biomedical literature to 'triangulate' evidence from different sources. AVAILABILITY AND IMPLEMENTATION: The EpiGraphDB platform is openly available at https://epigraphdb.org. Code for replicating case study results is available at https://github.com/MRCIEU/epigraphdb as Jupyter notebooks using the API, and https://mrcieu.github.io/epigraphdb-r using the R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi Liu 0058, Benjamin L. Elsworth, Pau Erola, Valeriia Haberland, Gibran Hemani, Matt Lyon, Jie Zheng 0006, Oliver Lloyd, Marina Vabistsevits, Tom R. Gaunt |
Bioinform. | 10 |
| 2021 | CScape-somatic: distinguishing driver and passenger point mutations in the cancer genomeabstractBioinformatics (2020) doi:10.1093/bioinformatics/btaa242 In the originally published article, a funding acknowledgement was missing. This should read: “Financial Support: The Integrative Epidemiology Unit is supported by the Medical Research Council (MC_UU_00011/4) and the University of Bristol, and we also acknowledge funding from the Cancer Research UK Integrative Cancer Epidemiology Programme (C18281/A19169).” instead of: “Financial Support: none declared.” These details have been corrected online. Mark F. Rogers, Tom R. Gaunt, Colin Campbell |
Bioinform. | 2 |
| 2021 | MendelVar: gene prioritization at GWAS loci using phenotypic enrichment of Mendelian disease genesabstractMOTIVATION: Gene prioritization at human GWAS loci is challenging due to linkage-disequilibrium and long-range gene regulatory mechanisms. However, identifying the causal gene is crucial to enable identification of potential drug targets and better understanding of molecular mechanisms. Mapping GWAS traits to known phenotypically relevant Mendelian disease genes near a locus is a promising approach to gene prioritization. RESULTS: We present MendelVar, a comprehensive tool that integrates knowledge from four databases on Mendelian disease genes with enrichment testing for a range of associated functional annotations such as Human Phenotype Ontology, Disease Ontology and variants from ClinVar. This open web-based platform enables users to strengthen the case for causal importance of phenotypically matched candidate genes at GWAS loci. We demonstrate the use of MendelVar in post-GWAS gene annotation for type 1 diabetes, type 2 diabetes, blood lipids and atopic dermatitis. AVAILABILITY AND IMPLEMENTATION: MendelVar is freely available at https://mendelvar.mrcieu.ac.uk. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. M. K. Sobczyk, Tom R. Gaunt, Lavinia Paternoster |
Bioinform. | 2 |
| 2020 | CScape-somatic: distinguishing driver and passenger point mutations in the cancer genomeabstractMOTIVATION: Next-generation sequencing technologies have accelerated the discovery of single nucleotide variants in the human genome, stimulating the development of predictors for classifying which of these variants are likely functional in disease, and which neutral. Recently, we proposed CScape, a method for discriminating between cancer driver mutations and presumed benign variants. For the neutral class, this method relied on benign germline variants found in the 1000 Genomes Project database. Discrimination could, therefore, be influenced by the distinction of germline versus somatic, rather than neutral versus disease driver. This motivates this article in which we consider predictive discrimination between recurrent and rare somatic single point mutations based solely on using cancer data, and the distinction between these two somatic classes and germline single point mutations. RESULTS: For somatic point mutations in coding and non-coding regions of the genome, we propose CScape-somatic, an integrative classifier for predictively discriminating between recurrent and rare variants in the human cancer genome. In this study, we use purely cancer genome data and investigate the distinction between minimal occurrence and significantly recurrent somatic single point mutations in the human cancer genome. We show that this type of predictive distinction can give novel insight, and may deliver more meaningful prediction in both coding and non-coding regions of the cancer genome. Tested on somatic mutations, CScape-somatic outperforms alternative methods, reaching 74% balanced accuracy in coding regions and 69% in non-coding regions, whereas even higher accuracy may be achieved using thresholds to isolate high-confidence predictions. AVAILABILITY AND IMPLEMENTATION: Predictions and software are available at http://CScape-somatic.biocompute.org.uk/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mark F. Rogers, Tom R. Gaunt, Colin Campbell |
Bioinform. | 2 |
| 2018 | FATHMM-XF: accurate prediction of pathogenic point mutations via extended featuresabstractSummary: We present FATHMM-XF, a method for predicting pathogenic point mutations in the human genome. Drawing on an extensive feature set, FATHMM-XF outperforms competitors on benchmark tests, particularly in non-coding regions where the majority of pathogenic mutations are likely to be found. Availability and implementation: The FATHMM-XF web server is available at http://fathmm.biocompute.org.uk/fathmm-xf/, and as tracks on the Genome Tolerance Browser: http://gtb.biocompute.org.uk. Predictions are provided for human genome version GRCh37/hg19. The data used for this project can be downloaded from: http://fathmm.biocompute.org.uk/fathmm-xf/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Mark F. Rogers, Hashem A. Shihab, Matthew E. Mort, David N. Cooper, Tom R. Gaunt, Colin Campbell |
Bioinform. | 5 |
| 2017 | HIPred: an integrative approach to predicting haploinsufficient genesabstractMOTIVATION: A major cause of autosomal dominant disease is haploinsufficiency, whereby a single copy of a gene is not sufficient to maintain the normal function of the gene. A large proportion of existing methods for predicting haploinsufficiency incorporate biological networks, e.g. protein-protein interaction networks that have recently been shown to introduce study bias. As a result, these methods tend to perform best on well-studied genes, but underperform on less studied genes. The advent of large genome sequencing consortia, such as the 1000 genomes project, NHLBI Exome Sequencing Project and the Exome Aggregation Consortium creates an urgent need for unbiased haploinsufficiency prediction methods. RESULTS: Here, we describe a machine learning approach, called HIPred, that integrates genomic and evolutionary information from ENSEMBL, with functional annotations from the Encyclopaedia of DNA Elements consortium and the NIH Roadmap Epigenomics Project to predict haploinsufficiency, without the study bias described earlier. We benchmark HIPred using several datasets and show that our unbiased method performs as well as, and in most cases, outperforms existing biased algorithms. AVAILABILITY AND IMPLEMENTATION: HIPred scores for all gene identifiers are available at: https://github.com/HAShihab/HIPred . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hashem A. Shihab, Mark F. Rogers, Colin Campbell, Tom R. Gaunt |
Bioinform. | 4 |
| 2017 | LD Hub: a centralized database and web interface to perform LD score regression that maximizes the potential of summary level GWAS data for SNP heritability and genetic correlation analysisabstractMOTIVATION: LD score regression is a reliable and efficient method of using genome-wide association study (GWAS) summary-level results data to estimate the SNP heritability of complex traits and diseases, partition this heritability into functional categories, and estimate the genetic correlation between different phenotypes. Because the method relies on summary level results data, LD score regression is computationally tractable even for very large sample sizes. However, publicly available GWAS summary-level data are typically stored in different databases and have different formats, making it difficult to apply LD score regression to estimate genetic correlations across many different traits simultaneously. RESULTS: In this manuscript, we describe LD Hub - a centralized database of summary-level GWAS results for 173 diseases/traits from different publicly available resources/consortia and a web interface that automates the LD score regression analysis pipeline. To demonstrate functionality and validate our software, we replicated previously reported LD score regression analyses of 49 traits/diseases using LD Hub; and estimated SNP heritability and the genetic correlation across the different phenotypes. We also present new results obtained by uploading a recent atopic dermatitis GWAS meta-analysis to examine the genetic correlation between the condition and other potentially related traits. In response to the growing availability of publicly accessible GWAS summary-level results data, our database and the accompanying web interface will ensure maximal uptake of the LD score regression methodology, provide a useful database for the public dissemination of GWAS results, and provide a method for easily screening hundreds of traits for overlapping genetic aetiologies. AVAILABILITY AND IMPLEMENTATION: The web interface and instructions for using LD Hub are available at http://ldsc.broadinstitute.org/ CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Jie Zheng 0006, Mesut A. Erzurumluoglu, Benjamin L. Elsworth, John P. Kemp, Laurence J. Howe, Philip C. Haycock, Gibran Hemani, Katherine Tansey, Charles Laurin, Early Genetics, Beate St. Pourcain, Nicole M. Warrington, Hilary K. Finucane, Alkes L. Price, Brendan K. Bulik-Sullivan, Verneri Anttila, Lavinia Paternoster, Tom R. Gaunt, David M. Evans 0002, Benjamin M. Neale |
Bioinform. | 19 |
| 2017 | HAPRAP: a haplotype-based iterative method for statistical fine mapping using GWAS summary statisticsabstractMOTIVATION: Fine mapping is a widely used approach for identifying the causal variant(s) at disease-associated loci. Standard methods (e.g. multiple regression) require individual level genotypes. Recent fine mapping methods using summary-level data require the pairwise correlation coefficients ([Formula: see text]) of the variants. However, haplotypes rather than pairwise [Formula: see text], are the true biological representation of linkage disequilibrium (LD) among multiple loci. In this article, we present an empirical iterative method, HAPlotype Regional Association analysis Program (HAPRAP), that enables fine mapping using summary statistics and haplotype information from an individual-level reference panel. RESULTS: Simulations with individual-level genotypes show that the results of HAPRAP and multiple regression are highly consistent. In simulation with summary-level data, we demonstrate that HAPRAP is less sensitive to poor LD estimates. In a parametric simulation using Genetic Investigation of ANthropometric Traits height data, HAPRAP performs well with a small training sample size (N < 2000) while other methods become suboptimal. Moreover, HAPRAP's performance is not affected substantially by single nucleotide polymorphisms (SNPs) with low minor allele frequencies. We applied the method to existing quantitative trait and binary outcome meta-analyses (human height, QTc interval and gallbladder disease); all previous reported association signals were replicated and two additional variants were independently associated with human height. Due to the growing availability of summary level data, the value of HAPRAP is likely to increase markedly for future analyses (e.g. functional prediction and identification of instruments for Mendelian randomization). AVAILABILITY AND IMPLEMENTATION: The HAPRAP package and documentation are available at http://apps.biocompute.org.uk/haprap/ CONTACT: : [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Jie Zheng 0006, Santiago Rodríguez, Charles Laurin, Denis Baird, Lea Trela-Larsen, Mesut A. Erzurumluoglu, Jon White, Claudia Giambartolomei, Delilah Zabaneh, Richard Morris, Meena Kumari, Juan P. Casas, Aroon D. Hingorani, David M. Evans 0002, Tom R. Gaunt, Ian N. M. Day |
Bioinform. | 16 |
| 2017 | An integrative approach to predicting the functional effects of small indels in non-coding regions of the human genomeabstractBACKGROUND: Small insertions and deletions (indels) have a significant influence in human disease and, in terms of frequency, they are second only to single nucleotide variants as pathogenic mutations. As the majority of mutations associated with complex traits are located outside the exome, it is crucial to investigate the potential pathogenic impact of indels in non-coding regions of the human genome. RESULTS: We present FATHMM-indel, an integrative approach to predict the functional effect, pathogenic or neutral, of indels in non-coding regions of the human genome. Our method exploits various genomic annotations in addition to sequence data. When validated on benchmark data, FATHMM-indel significantly outperforms CADD and GAVIN, state of the art models in assessing the pathogenic impact of non-coding variants. FATHMM-indel is available via a web server at indels.biocompute.org.uk. CONCLUSIONS: FATHMM-indel can accurately predict the functional impact and prioritise small indels throughout the whole non-coding genome. Michael Ferlaino, Mark F. Rogers, Hashem A. Shihab, Matthew E. Mort, David N. Cooper, Tom R. Gaunt, Colin Campbell |
BMC Bioinform. | 6 |
| 2017 | GTB - an online genome tolerance browserabstractBACKGROUND: Accurate methods capable of predicting the impact of single nucleotide variants (SNVs) are assuming ever increasing importance. There exists a plethora of in silico algorithms designed to help identify and prioritize SNVs across the human genome for further investigation. However, no tool exists to visualize the predicted tolerance of the genome to mutation, or the similarities between these methods. RESULTS: We present the Genome Tolerance Browser (GTB, http://gtb.biocompute.org.uk ): an online genome browser for visualizing the predicted tolerance of the genome to mutation. The server summarizes several in silico prediction algorithms and conservation scores: including 13 genome-wide prediction algorithms and conservation scores, 12 non-synonymous prediction algorithms and four cancer-specific algorithms. CONCLUSION: The GTB enables users to visualize the similarities and differences between several prediction algorithms and to upload their own data as additional tracks; thereby facilitating the rapid identification of potential regions of interest. Hashem A. Shihab, Mark F. Rogers, Michael Ferlaino, Colin Campbell, Tom R. Gaunt |
BMC Bioinform. | 5 |
| 2015 | Sequential data selection for predicting the pathogenic effects of sequence variationabstractRecent improvements in sequencing technologies provide unprecedented opportunities to investigate the role of genetic variation in human disease. In previous work we have proposed a machine learning approach to predicting whether single nucleotide variants (SNVs) are functional or neutral in human disease. Many data sources from the Encyclopaedia of DNA Elements (ENCODE) may be relevant to this problem. To integrate these data sources, we applied integrative multiple kernel learning (MKL) that weights each source according to its relevance. Using an MKL optimization that yields sparse weights, we were able to eliminate the least informative data sources from our model. However, when selecting from a wide assortment of data sources, we have found that MKL may not be an efficient method for eliminating uninformative sources. Many data sources related to the human genome are incomplete: this can reduce dramatically the data available for training and the proportion of novel predictions that exploit all relevant sources. Here we introduce a greedy sequential selection method that assesses data sources in a structured fashion prior to MKL weight optimization. This method allows us to eliminate a majority of uninformative data sources prior to assigning kernel weights. When we use this method with our coding-region predictor, we select just five kernels for our final model, yielding increased accuracy over our previous model. In addition, by reducing the amount of data required for novel predictions, we are able to increase by five fold our model's coverage for new predictions. Mark F. Rogers, Colin Campbell, Hashem A. Shihab, Tom R. Gaunt, Matthew E. Mort, David N. Cooper |
BIBM | 4 |
| 2015 | An integrative approach to predicting the functional effects of non-coding and coding sequence variationabstractAbstract Motivation: Technological advances have enabled the identification of an increasingly large spectrum of single nucleotide variants within the human genome, many of which may be associated with monogenic disease or complex traits. Here, we propose an integrative approach, named FATHMM-MKL, to predict the functional consequences of both coding and non-coding sequence variants. Our method utilizes various genomic annotations, which have recently become available, and learns to weight the significance of each component annotation source. Results: We show that our method outperforms current state-of-the-art algorithms, CADD and GWAVA, when predicting the functional consequences of non-coding variants. In addition, FATHMM-MKL is comparable to the best of these algorithms when predicting the impact of coding variants. The method includes a confidence measure to rank order predictions. Availability and implementation: The FATHMM-MKL webserver is available at: http://fathmm.biocompute.org.uk Contact: [email protected] or [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Hashem A. Shihab, Mark F. Rogers, Julian Gough, Matthew E. Mort, David N. Cooper, Ian N. M. Day, Tom R. Gaunt, Colin Campbell |
Bioinform. | 7 |
| 2015 | Texture classification using feature selection and kernel-based techniques
Carlos Fernandez-Lozano, José Antonio Seoane Fernández, Marcos Gestal, Tom R. Gaunt, Julián Dorado, Colin Campbell |
Soft Comput. | 4 |
| 2014 | A Random Forest proximity matrix as a new measure for gene annotation
José Antonio Seoane Fernández, Ian N. M. Day, Juan P. Casas, Colin Campbell, Tom R. Gaunt |
ESANN | 5 |
| 2014 | A pathway-based data integration framework for prediction of disease progressionabstractMOTIVATION: Within medical research there is an increasing trend toward deriving multiple types of data from the same individual. The most effective prognostic prediction methods should use all available data, as this maximizes the amount of information used. In this article, we consider a variety of learning strategies to boost prediction performance based on the use of all available data. IMPLEMENTATION: We consider data integration via the use of multiple kernel learning supervised learning methods. We propose a scheme in which feature selection by statistical score is performed separately per data type and by pathway membership. We further consider the introduction of a confidence measure for the class assignment, both to remove some ambiguously labeled datapoints from the training data and to implement a cautious classifier that only makes predictions when the associated confidence is high. RESULTS: We use the METABRIC dataset for breast cancer, with prediction of survival at 2000 days from diagnosis. Predictive accuracy is improved by using kernels that exclusively use those genes, as features, which are known members of particular pathways. We show that yet further improvements can be made by using a range of additional kernels based on clinical covariates such as Estrogen Receptor (ER) status. Using this range of measures to improve prediction performance, we show that the test accuracy on new instances is nearly 80%, though predictions are only made on 69.2% of the patient cohort. AVAILABILITY: https://github.com/jseoane/FSMKL CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. José Antonio Seoane Fernández, Ian N. M. Day, Tom R. Gaunt, Colin Campbell |
Bioinform. | 3 |
| 2014 | Canonical Correlation Analysis for Gene-Based Pleiotropy DiscoveryabstractGenome-wide association studies have identified a wealth of genetic variants involved in complex traits and multifactorial diseases. There is now considerable interest in testing variants for association with multiple phenotypes (pleiotropy) and for testing multiple variants for association with a single phenotype (gene-based association tests). Such approaches can increase statistical power by combining evidence for association over multiple phenotypes or genetic variants respectively. Canonical Correlation Analysis (CCA) measures the correlation between two sets of multidimensional variables, and thus offers the potential to combine these two approaches. To apply CCA, we must restrict the number of attributes relative to the number of samples. Hence we consider modules of genetic variation that can comprise a gene, a pathway or another biologically relevant grouping, and/or a set of phenotypes. In order to do this, we use an attribute selection strategy based on a binary genetic algorithm. Applied to a UK-based prospective cohort study of 4286 women (the British Women's Heart and Health Study), we find improved statistical power in the detection of previously reported genetic associations, and identify a number of novel pleiotropic associations between genetic variants and phenotypes. New discoveries include gene-based association of NSF with triglyceride levels and several genes (ACSM3, ERI2, IL18RAP, IL23RAP and NRG1) with left ventricular hypertrophy phenotypes. In multiple-phenotype analyses we find association of NRG1 with left ventricular hypertrophy phenotypes, fibrinogen and urea and pleiotropic relationships of F7 and F10 with Factor VII, Factor IX and cholesterol levels. José Antonio Seoane Fernández, Colin Campbell, Ian N. M. Day, Juan P. Casas, Tom R. Gaunt |
PLoS Comput. Biol. | 5 |
| 2013 | Predicting the functional consequences of cancer-associated amino acid substitutionsabstractMOTIVATION: The number of missense mutations being identified in cancer genomes has greatly increased as a consequence of technological advances and the reduced cost of whole-genome/whole-exome sequencing methods. However, a high proportion of the amino acid substitutions detected in cancer genomes have little or no effect on tumour progression (passenger mutations). Therefore, accurate automated methods capable of discriminating between driver (cancer-promoting) and passenger mutations are becoming increasingly important. In our previous work, we developed the Functional Analysis through Hidden Markov Models (FATHMM) software and, using a model weighted for inherited disease mutations, observed improved performances over alternative computational prediction algorithms. Here, we describe an adaptation of our original algorithm that incorporates a cancer-specific model to potentiate the functional analysis of driver mutations. RESULTS: The performance of our algorithm was evaluated using two separate benchmarks. In our analysis, we observed improved performances when distinguishing between driver mutations and other germ line variants (both disease-causing and putatively neutral mutations). In addition, when discriminating between somatic driver and passenger mutations, we observed performances comparable with the leading computational prediction algorithms: SPF-Cancer and TransFIC. AVAILABILITY AND IMPLEMENTATION: A web-based implementation of our cancer-specific model, including a downloadable stand-alone package, is available at http://fathmm.biocompute.org.uk. Hashem A. Shihab, Julian Gough, David N. Cooper, Ian N. M. Day, Tom R. Gaunt |
Bioinform. | 5 |
| 2007 | Cubic exact solutions for the estimation of pairwise haplotype frequencies: implications for linkage disequilibrium analyses and a web tool 'CubeX'abstractBACKGROUND: The frequency of a haplotype comprising one allele at each of two loci can be expressed as a cubic equation (the 'Hill equation'), the solution of which gives that frequency. Most haplotype and linkage disequilibrium analysis programs use iteration-based algorithms which substitute an estimate of haplotype frequency into the equation, producing a new estimate which is repeatedly fed back into the equation until the values converge to a maximum likelihood estimate (expectation-maximisation). RESULTS: We present a program, "CubeX", which calculates the biologically possible exact solution(s) and provides estimated haplotype frequencies, D', r2 and chi2 values for each. CubeX provides a "complete" analysis of haplotype frequencies and linkage disequilibrium for a pair of biallelic markers under situations where sampling variation and genotyping errors distort sample Hardy-Weinberg equilibrium, potentially causing more than one biologically possible solution. We also present an analysis of simulations and real data using the algebraically exact solution, which indicates that under perfect sample Hardy-Weinberg equilibrium there is only one biologically possible solution, but that under other conditions there may be more. CONCLUSION: Our analyses demonstrate that lower allele frequencies, lower sample numbers, population stratification and a possible |D'| value of 1 are particularly susceptible to distortion of sample Hardy-Weinberg equilibrium, which has significant implications for calculation of linkage disequilibrium in small sample sizes (eg HapMap) and rarer alleles (eg paucimorphisms, q < 0.05) that may have particular disease relevance and require improved approaches for meaningful evaluation. Tom R. Gaunt, Santiago Rodríguez, Ian N. M. Day |
BMC Bioinform. | 1 |
| 2006 | MIDAS: software for analysis and visualisation of interallelic disequilibrium between multiallelic markersabstractBACKGROUND: Various software tools are available for the display of pairwise linkage disequilibrium across multiple single nucleotide polymorphisms. The HapMap project also presents these graphics within their website. However, these approaches are limited in their use of data from multiallelic markers and provide limited information in a graphical form. RESULTS: We have developed a software package (MIDAS - Multiallelic Interallelic Disequilibrium Analysis Software) for the estimation and graphical display of interallelic linkage disequilibrium. Linkage disequilibrium is analysed for each allelic combination (of one allele from each of two loci), between all pairwise combinations of any type of multiallelic loci in a contig (or any set) of many loci (including single nucleotide polymorphisms, microsatellites, minisatellites and haplotypes). Data are presented graphically in a novel and informative way, and can also be exported in tabular form for other analyses. This approach facilitates visualisation of patterns of linkage disequilibrium across genomic regions, analysis of the relationships between different alleles of multiallelic markers and inferences about patterns of evolution and selection. CONCLUSION: MIDAS is a linkage disequilibrium analysis program with a comprehensive graphical user interface providing novel views of patterns of linkage disequilibrium between all types of multiallelic and biallelic markers. Tom R. Gaunt, Santiago Rodríguez, Carlos Zapata, Ian N. M. Day |
BMC Bioinform. | 1 |