Ross D. King

dblp:k/RossDKing · DBLP profile ↗
← Back
68ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0001-7208-4387ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 33 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-authorTheory of computation · 8 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Drug response profile-based machine learning enables strategic cell line and compound selection for drug development
abstract
MOTIVATION: Early-stage drug discovery relies on testing compounds across a limited set of cell lines, making it challenging to capture biological diversity while maintaining experimental efficiency. Current predictive approaches for identifying responsive cell lines often depend on high-dimensional omics data, which can be costly and difficult to interpret. We therefore evaluated whether drug-response panel (DRP) descriptors, which capture sensitivity profiles to a reference set of compounds, can provide an efficient and informative alternative for modelling drug response in cell lines. RESULTS: Using gradient boosting models across GDSC and CCLE datasets, DRP descriptors consistently outperformed mRNA expression features in predicting drug sensitivity (-log10(IC50)), although performance varied across compounds. DRP-guided cell line selection enabled downstream omics-based modelling that recovered known MAPK-associated sensitivity signatures and identified potential biomarkers for MEK1/2 and BTK/MNK inhibitors. Extending this framework, we demonstrated its utility in compound prioritisation by distinguishing between tumourigenic MCF7 and non-tumourigenic MCF10A cells, successfully identifying compounds with selective activity. Together, these results show that DRP-based representations, derived from compact screening panels, support efficient cell line selection, biomarker discovery, and compound prioritisation in early-stage drug development. AVAILABILITY: Code and data uploaded to https://github.com/abbiAR/-Strategic-Cell-Line-and-Compound-Selection-Using-Drug-Response-Profiles.
Abbi Abdel-Rehim, Emma Tate, Larisa N. Soldatova, Ross D. King
Bioinform.4
2026 Automated Scientific Discovery: From Equation Discovery to Autonomous Discovery Systems
abstract
Abstract The paper surveys automated scientific discovery, from equation discovery and symbolic regression to autonomous discovery systems and agents. It discusses the individual approaches from a "big picture" perspective and in context, but also discusses open issues and recent topics like the various roles of deep neural networks in this area, aiding in the discovery of human-interpretable knowledge. Further, we will present closed-loop scientific discovery systems, starting with the pioneering work on the Adam system up to current efforts in fields from material science to astronomy. Finally, we will elaborate on autonomy from a machine learning perspective, but also in analogy to the autonomy levels in autonomous driving. The maximal level, level five, is defined to require no human intervention at all in the production of scientific knowledge. Achieving this is one step towards solving the Nobel Turing Grand Challenge to develop AI Scientists: AI systems capable of making Nobel-quality scientific discoveries highly autonomously at a level comparable, and possibly superior, to the best human scientists by 2050.
Stefan Kramer 0001, Mattia Cerrato, Jannis Brugger, Saso Dzeroski, Ross D. King
Mach. Learn.5
2025 Ontology-based box embeddings and knowledge graphs for predicting phenotypic traits in Saccharomyces cerevisiae
abstract
We present a method that uses graph neural networks (GNNs) to predict and interpret digenic deletion fitness in the yeast Saccharomyces cerevisiae from a knowledge graph (KG) with ontology-based box embeddings. We construct the KG from community databases using terms defined in several ontologies. From the class hierarchies in the ontologies, box embeddings are learnt as low dimensional representations of the nodes in the graph, which are used together with GNNs to predict cell growth for digenic deletions from the KG. With this we show that high level qualitative information can be used to predict experimental data. Prediction performance was improved when using box embeddings of ontologies to represent the nodes in the graph, compared to learning features specific for this task. This suggests that class hierarchies in ontologies contain useful information about the domains, which can be extracted in the training of the box embeddings. We also demonstrate that our model can generalise beyond the task it was trained for by evaluating it on higher order gene deletions. Additionally, we apply model interpretability techniques to identify co-occurring edges critical for fitness. Our findings are further validated by a biological experiment that reveals an association between inositol utilization and osmotic stress resistance, emphasising the model’s potential to guide scientific discovery.
Filip Kronström, Daniel Brunnsåker, Ievgeniia A. Tiukova, Ross D. King
NeSy4
2024 Random Forests for Heteroscedastic Data
Hugo Bellamy, Ross D. King
DS (2)2
2024 Interpreting protein abundance in Saccharomyces cerevisiae through relational learning
abstract
MOTIVATION: Proteomic profiles reflect the functional readout of the physiological state of an organism. An increased understanding of what controls and defines protein abundances is of high scientific interest. Saccharomyces cerevisiae is a well-studied model organism, and there is a large amount of structured knowledge on yeast systems biology in databases such as the Saccharomyces Genome Database, and highly curated genome-scale metabolic models like Yeast8. These datasets, the result of decades of experiments, are abundant in information, and adhere to semantically meaningful ontologies. RESULTS: By representing this knowledge in an expressive Datalog database we generated data descriptors using relational learning that, when combined with supervised machine learning, enables us to predict protein abundances in an explainable manner. We learnt predictive relationships between protein abundances, function and phenotype; such as α-amino acid accumulations and deviations in chronological lifespan. We further demonstrate the power of this methodology on the proteins His4 and Ilv2, connecting qualitative biological concepts to quantified abundances. AVAILABILITY AND IMPLEMENTATION: All data and processing scripts are available at the following Github repository: https://github.com/DanielBrunnsaker/ProtPredict.
Daniel Brunnsåker, Filip Kronström, Ievgeniia A. Tiukova, Ross D. King
Bioinform.4
2024 Extrapolation is not the same as interpolation
abstract
Abstract We propose a new machine learning formulation designed specifically for extrapolation. The textbook way to apply machine learning to drug design is to learn a univariate function that when a drug (structure) is input, the function outputs a real number (the activity): f(drug) $$\rightarrow$$ → activity. However, experience in real-world drug design suggests that this formulation of the drug design problem is not quite correct. Specifically, what one is really interested in is extrapolation: predicting the activity of new drugs with higher activity than any existing ones. Our new formulation for extrapolation is based on learning a bivariate function that predicts the difference in activities of two drugs F(drug1, drug2) $$\rightarrow$$ → difference in activity, followed by the use of ranking algorithms. This formulation is general and agnostic, suitable for finding samples with target values beyond the target value range of the training set. We applied the formulation to work with support vector machines , random forests , and Gradient Boosting Machines . We compared the formulation with standard regression on thousands of drug design datasets, gene expression datasets and material property datasets. The test set extrapolation metric was the identification of examples with greater values than the training set, and top-performing examples (within the top 10% of the whole dataset). On this metric our pairwise formulation vastly outperformed standard regression. Its proposed variations also showed a consistent outperformance. Its application in the stock selection problem further confirmed the advantage of this pairwise formulation.
Ross D. King
Mach. Learn.2
2023 LGEM+: A First-Order Logic Framework for Automated Improvement of Metabolic Network Models Through Abduction
abstract
Abstract Scientific discovery in biology is difficult due to the complexity of the systems involved and the expense of obtaining high quality experimental data. Automated techniques are a promising way to make scientific discoveries at the scale and pace required to model large biological systems. A key problem for 21st century biology is to build a computational model of the eukaryotic cell. The yeast Saccharomyces cerevisiae is the best understood eukaryote, and genome-scale metabolic models (GEMs) are rich sources of background knowledge that we can use as a basis for automated inference and investigation. We present LGEM+, a system for automated abductive improvement of GEMs consisting of: a compartmentalised first-order logic framework for describing biochemical pathways (using curated GEMs as the expert knowledge source); and a two-stage hypothesis abduction procedure. We demonstrate that deductive inference on logical theories created using LGEM+, using the automated theorem prover iProver, can predict growth/no-growth of S. cerevisiae strains in minimal media. LGEM+ proposed 2094 unique candidate hypotheses for model improvement. We assess the value of the generated hypotheses using two criteria: (a) genome-wide single-gene essentiality prediction, and (b) constraint of flux-balance analysis (FBA) simulations. For (b) we developed an algorithm to integrate FBA with the logic model. We rank and filter the hypotheses using these assessments. We intend to test these hypotheses using the robot scientist Genesis, which is based around chemostat cultivation and high-throughput metabolomics.
Alexander H. Gower, Konstantin Korovin, Daniel Brunnsåker, Ievgeniia A. Tiukova, Ross D. King
DS5
2023 RIMBO - An Ontology for Model Revision Databases
abstract
Abstract The use of computational models is growing throughout most scientific domains. The increased complexity of such models, as well as the increased automation of scientific research, imply that model revisions need to be systematically recorded. We present RIMBO (Revisions for Improvements of Models in Biology Ontology), which describes the changes made to computational biology models. The ontology is intended as the foundation of a database containing and describing iterative improvements to models. By recording high level information, such as modelled phenomena, and model type, using controlled vocabularies from widely used ontologies, the same database can be used for different model types. The database aims to describe the evolution of models by recording chains of changes to them. To make this evolution transparent, emphasise has been put on recording the reasons, and descriptions, of the changes. We demonstrate the usefulness of a database based on this ontology by modelling the update from version 8.4.1 to 8.4.2 of the genome-scale metabolic model Yeast8, a modification proposed by an abduction algorithm, as well as thousands of simulated revisions. This results in a database demonstrating that revisions can successfully be modelled in a semantically meaningful and storage efficient way. We believe such a database is necessary for performing automated model improvement at scale in systems biology, as well as being a useful tool to increase the openness and traceability for model development. With minor modifications the ontology can also be used in other scientific domains. The ontology is made available at https://github.com/filipkro/rimbo and will be continually updated.
Filip Kronström, Alexander H. Gower, Ievgeniia A. Tiukova, Ross D. King
DS4
2023 Extrapolation is Not the Same as Interpolation
abstract
Abstract We propose a new machine learning formulation designed specifically for extrapolation. The textbook way to apply machine learning to drug design is to learn a univariate function that when a drug (structure) is input, the function outputs a real number (the activity): F(drug) → activity. The PubMed server lists around twenty thousand papers doing this. However, experience in real-world drug design suggests that this formulation of the drug design problem is not quite correct. Specifically, what one is really interested in is extrapolation: predicting the activity of new drugs with higher activity than any existing ones. Our new formulation for extrapolation is based around learning a bivariate function that predicts the difference in activities of two drugs: F(drug1, drug2) → signed difference in activity. This formulation is general and potentially suitable for problems to find samples with target values beyond the target value range of the training set. We applied the formulation to work with support vector machines (SVMs), random forests (RFs), and Gradient Boosting Machines (XGBs). We compared the formulation with standard regression on thousands of drug design datasets, and hundreds of gene expression datasets. The test set extrapolation metrics use the concept of classification metrics to count the identification of extraordinary examples (with greater values than the training set), and top-performing examples (within the top 10% of the whole dataset). On these metrics our pairwise formulation vastly outperformed standard regression for SVMs, RFs, and XGBs. We expect this success to extrapolate to other extrapolation problems.
Ross D. King
DS2
2023 Protein-ligand binding affinity prediction exploiting sequence constituent homology
abstract
MOTIVATION: Molecular docking is a commonly used approach for estimating binding conformations and their resultant binding affinities. Machine learning has been successfully deployed to enhance such affinity estimations. Many methods of varying complexity have been developed making use of some or all the spatial and categorical information available in these structures. The evaluation of such methods has mainly been carried out using datasets from PDBbind. Particularly the Comparative Assessment of Scoring Functions (CASF) 2007, 2013, and 2016 datasets with dedicated test sets. This work demonstrates that only a small number of simple descriptors is necessary to efficiently estimate binding affinity for these complexes without the need to know the exact binding conformation of a ligand. RESULTS: The developed approach of using a small number of ligand and protein descriptors in conjunction with gradient boosting trees demonstrates high performance on the CASF datasets. This includes the commonly used benchmark CASF2016 where it appears to perform better than any other approach. This methodology is also useful for datasets where the spatial relationship between the ligand and protein is unknown as demonstrated using a large ChEMBL-derived dataset. AVAILABILITY AND IMPLEMENTATION: Code and data uploaded to https://github.com/abbiAR/PLBAffinity.
Abbi Abdel-Rehim, Oghenejokpeme I. Orhobor, Hang Lou, Ross D. King
Bioinform.5
2023 Imbalanced regression using regressor-classifier ensembles
abstract
Abstract We present an extension to the federated ensemble regression using classification algorithm, an ensemble learning algorithm for regression problems which leverages the distribution of the samples in a learning set to achieve improved performance. We evaluated the extension using four classifiers and four regressors, two discretizers, and 119 responses from a wide variety of datasets in different domains. Additionally, we compared our algorithm to two resampling methods aimed at addressing imbalanced datasets. Our results show that the proposed extension is highly unlikely to perform worse than the base case, and on average outperforms the two resampling methods with significant differences in performance.
Oghenejokpeme I. Orhobor, Nastasiya F. Grinberg, Larisa N. Soldatova, Ross D. King
Mach. Learn.4
2022 Improved prediction of gene expression through integrating cell signalling models with machine learning
abstract
BACKGROUND: A key problem in bioinformatics is that of predicting gene expression levels. There are two broad approaches: use of mechanistic models that aim to directly simulate the underlying biology, and use of machine learning (ML) to empirically predict expression levels from descriptors of the experiments. There are advantages and disadvantages to both approaches: mechanistic models more directly reflect the underlying biological causation, but do not directly utilize the available empirical data; while ML methods do not fully utilize existing biological knowledge. RESULTS: Here, we investigate overcoming these disadvantages by integrating mechanistic cell signalling models with ML. Our approach to integration is to augment ML with similarity features (attributes) computed from cell signalling models. Seven sets of different similarity feature were generated using graph theory. Each set of features was in turn used to learn multi-target regression models. All the features have significantly improved accuracy over the baseline model - without the similarity features. Finally, the seven multi-target regression models were stacked together to form an overall prediction model that was significantly better than the baseline on 95% of genes on an independent test set. The similarity features enable this stacking model to provide interpretable knowledge about cancer, e.g. the role of ERBB3 in the MCF7 breast cancer cell line. CONCLUSION: Integrating mechanistic models as graphs helps to both improve the predictive results of machine learning models, and to provide biological knowledge about genes that can help in building state-of-the-art mechanistic models.
Nada Al taweraqi, Ross D. King
BMC Bioinform.2
2020 Generating Explainable and Effective Data Descriptors Using Relational Learning: Application to Cancer Biology
abstract
Abstract The key to success in machine learning is the use of effective data representations. The success of deep neural networks (DNNs) is based on their ability to utilize multiple neural network layers, and big data, to learn how to convert simple input representations into richer internal representations that are effective for learning. However, these internal representations are sub-symbolic and difficult to explain. In many scientific problems explainable models are required, and the input data is semantically complex and unsuitable for DNNs. This is true in the fundamental problem of understanding the mechanism of cancer drugs, which requires complex background knowledge about the functions of genes/proteins, their cells, and the molecular structure of the drugs. This background knowledge cannot be compactly expressed propositionally, and requires at least the expressive power of Datalog. Here we demonstrate the use of relational learning to generate new data descriptors in such semantically complex background knowledge. These new descriptors are effective: adding them to standard propositional learning methods significantly improves prediction accuracy. They are also explainable, and add to our understanding of cancer. Our approach can readily be expanded to include other complex forms of background knowledge, and combines the generality of relational learning with the efficiency of standard propositional learning.
Oghenejokpeme I. Orhobor, Joseph French, Larisa N. Soldatova, Ross D. King
DS4
2020 Federated Ensemble Regression Using Classification
abstract
Abstract Ensemble learning has been shown to significantly improve predictive accuracy in a variety of machine learning problems. For a given predictive task, the goal of ensemble learning is to improve predictive accuracy by combining the predictive power of multiple models. In this paper, we present an ensemble learning algorithm for regression problems which leverages the distribution of the samples in a learning set to achieve improved performance. We apply the proposed algorithm to a problem in precision medicine where the goal is to predict drug perturbation effects on genes in cancer cell lines. The proposed approach significantly outperforms the base case.
Oghenejokpeme I. Orhobor, Larisa N. Soldatova, Ross D. King
DS3
2020 An evaluation of machine-learning for predicting phenotype: studies in yeast, rice, and wheat
abstract
In phenotype prediction the physical characteristics of an organism are predicted from knowledge of its genotype and environment. Such studies, often called genome-wide association studies, are of the highest societal importance, as they are of central importance to medicine, crop-breeding, etc. We investigated three phenotype prediction problems: one simple and clean (yeast), and the other two complex and real-world (rice and wheat). We compared standard machine learning methods; elastic net, ridge regression, lasso regression, random forest, gradient boosting machines (GBM), and support vector machines (SVM), with two state-of-the-art classical statistical genetics methods; genomic BLUP and a two-step sequential method based on linear regression. Additionally, using the clean yeast data, we investigated how performance varied with the complexity of the biological mechanism, the amount of observational noise, the number of examples, the amount of missing data, and the use of different data representations. We found that for almost all the phenotypes considered, standard machine learning methods outperformed the methods from classical statistical genetics. On the yeast problem, the most successful method was GBM, followed by lasso regression, and the two statistical genetics methods; with greater mechanistic complexity GBM was best, while in simpler cases lasso was superior. In the wheat and rice studies the best two methods were SVM and BLUP. The most robust method in the presence of noise, missing data, etc. was random forests. The classical statistical genetics method of genomic BLUP was found to perform well on problems where there was population structure. This suggests that standard machine learning methods need to be refined to include population structure information when this is present. We conclude that the application of machine learning methods to phenotype prediction problems holds great promise, but that determining which methods is likely to perform well on any given problem is elusive and non-trivial.
Nastasiya F. Grinberg, Oghenejokpeme I. Orhobor, Ross D. King
Mach. Learn.3
2020 Predicting rice phenotypes with meta and multi-target learning
abstract
Abstract The features in some machine learning datasets can naturally be divided into groups. This is the case with genomic data, where features can be grouped by chromosome. In many applications it is common for these groupings to be ignored, as interactions may exist between features belonging to different groups. However, including a group that does not influence a response introduces noise when fitting a model, leading to suboptimal predictive accuracy. Here we present two general frameworks for the generation and combination of meta-features when feature groupings are present. Furthermore, we make comparisons to multi-target learning, given that one is typically interested in predicting multiple phenotypes. We evaluated the frameworks and multi-target learning approaches on a genomic rice dataset where the regression task is to predict plant phenotype. Our results demonstrate that there are use cases for both the meta and multi-target approaches, given that overall, they significantly outperform the base case.
Oghenejokpeme I. Orhobor, Nickolai N. Alexandrov, Ross D. King
Mach. Learn.3
2019 Using Prior Knowledge to Facilitate Computational Reading of Arabic Calligraphy
Seetah ALSalamah, Riza Theresa Batista-Navarro, Ross D. King
IDEAL (2)3
2018 Predicting Rice Phenotypes with Meta-learning
Oghenejokpeme I. Orhobor, Nickolai N. Alexandrov, Ross D. King
DS3
2018 Large-Scale Assessment of Deep Relational Machines
Tirtharaj Dash, Ashwin Srinivasan 0001, Lovekesh Vig, Oghenejokpeme I. Orhobor, Ross D. King
ILP5
2018 Meta-QSAR: a large-scale application of meta-learning to drug design and discovery
abstract
We investigate the learning of quantitative structure activity relationships (QSARs) as a case-study of meta-learning. This application area is of the highest societal importance, as it is a key step in the development of new medicines. The standard QSAR learning problem is: given a target (usually a protein) and a set of chemical compounds (small molecules) with associated bioactivities (e.g. inhibition of the target), learn a predictive mapping from molecular representation to activity. Although almost every type of machine learning method has been applied to QSAR learning there is no agreed single best way of learning QSARs, and therefore the problem area is well-suited to meta-learning. We first carried out the most comprehensive ever comparison of machine learning methods for QSAR learning: 18 regression methods, 3 molecular representations, applied to more than 2700 QSAR problems. (These results have been made publicly available on OpenML and represent a valuable resource for testing novel meta-learning methods.) We then investigated the utility of algorithm selection for QSAR problems. We found that this meta-learning approach outperformed the best individual QSAR learning method (random forests using a molecular fingerprint representation) by up to 13%, on average. We conclude that meta-learning outperforms base-learning methods for QSAR learning, and as this investigation is one of the most extensive ever comparisons of base and meta-learning methods ever made, it provides evidence for the general effectiveness of meta-learning over base-learning.
Iván Olier, Noureddin Sadawi, G. Richard J. Bickerton, Joaquin Vanschoren, Crina Grosan, Larisa N. Soldatova, Ross D. King
Mach. Learn.7
2014 Predicting the Geographical Origin of Music
abstract
Traditional research into the arts has almost always been based around the subjective judgment of human critics. The use of data mining tools to understand art has great promise as it is objective and operational. We investigate the distribution of music from around the world: geographical ethnomusicology. We cast the problem as training a machine learning program to predict the geographical origin of pieces of music. This is a technically interesting problem as it has features of both classification and regression, and because of the spherical geometry of the surface of the Earth. Because of these characteristics of the representation of geographical positions, most standard classification/regression methods cannot be directly used. Two applicable methods are K-Nearest Neighbors and Random forest regression, which are robust to the non-standard structure of data. We also investigated improving performance through use of bagging. We collected 1,142 pieces of music from 73 countries/areas, and described them using 2 different sets of standard audio descriptors using MARSYAS. 10-fold cross validation was used in all experiments. The experimental results indicate that Random forest regression produces significantly better results than KNN, and the use of bagging improves the performance of KNN. The best performing algorithm achieved a mean great circle distance error of 3,113 km.
Claire Q, Ross D. King
ICDM3
2014 EXACT2: the semantics of biomedical protocols
abstract
BACKGROUND: The reliability and reproducibility of experimental procedures is a cornerstone of scientific practice. There is a pressing technological need for the better representation of biomedical protocols to enable other agents (human or machine) to better reproduce results. A framework that ensures that all information required for the replication of experimental protocols is essential to achieve reproducibility. To construct EXACT2 we manually inspected hundreds of published and commercial biomedical protocols from several areas of biomedicine. After establishing a clear pattern for extracting the required information we utilized text-mining tools to translate the protocols into a machine amenable format. We have verified the utility of EXACT2 through the successful processing of previously 'unseen' (not used for the construction of EXACT2)protocols. METHODS: We have developed the ontology EXACT2 (EXperimental ACTions) that is designed to capture the full semantics of biomedical protocols required for their reproducibility. RESULTS: The paper reports on a fundamentally new version EXACT2 that supports the semantically-defined representation of biomedical protocols. The ability of EXACT2 to capture the semantics of biomedical procedures was verified through a text mining use case. In this EXACT2 is used as a reference model for text mining tools to identify terms pertinent to experimental actions, and their properties, in biomedical protocols expressed in natural language. An EXACT2-based framework for the translation of biomedical protocols to a machine amenable format is proposed. CONCLUSIONS: The EXACT2 ontology is sufficient to record, in a machine processable form, the essential information about biomedical protocols. EXACT2 defines explicit semantics of experimental actions, and can be used by various computer applications. It can serve as a reference model for for the translation of biomedical protocols in natural language into a semantically-defined format.
Larisa N. Soldatova, Daniel Nadis, Ross D. King, Piyali S. Basu, Emma Haddi, Véronique Baumlé, Nigel J. Saunders, Wolfgang Marwan, Brian B. Rudkin
BMC Bioinform.3
2012 Topic Models with Relational Features for Drug Design
Tanveer A. Faruquie, Ashwin Srinivasan 0001, Ross D. King
ILP3
2010 Logic-Based Steady-State Analysis and Revision of Metabolic Networks with Inhibition
abstract
This paper presents a qualitative logic-based method for the steady-state analysis and revision of metabolic networks with inhibition. The approach is able to automatically revise an initial metabolic model - through the addition and removal of whole reactions or individual substrates, products and inhibitors - in order to ensure the existence of a steady-state behaviour consistent with a set of experimental observations. We show how this can be done in a nonmonotonic logic programming setting and discuss the challenges that arise when metabolic cycles or mutual inhibitions occur in the underlying network.
Oliver Ray, Ken E. Whelan, Ross D. King
CISIS3
2009 A Nonmonotonic Logical Approach for Modelling and Revising Metabolic Networks
abstract
This paper describes a new logic-based approach for representing and reasoning about metabolic networks.First it shows how biological pathways can be elegantly represented in a logic programming formalism able to model full chemical reactions with substrates and products in different cell compartments, and which are catalysed by iso-enzymes or enzyme-complexes that are subject to inhibitory feedbacks.Then it shows how a nonmonotonic reasoning system called XHAIL can be used as a practical method for learning and revising such metabolic networks from observational data. Preliminary results are described in which the approach is validated on a state-of-the-art model of aromatic amino acid biosynthesis.
Oliver Ray, Ken E. Whelan, Ross D. King
CISIS3
2009 Automatic Revision of Metabolic Networks through Logical Analysis of Experimental Data
Oliver Ray, Ken E. Whelan, Ross D. King
ILP3
2009 An investigation into the population abundance distribution of mRNAs, proteins, and metabolites in biological systems
abstract
MOTIVATION: Distribution analysis is one of the most basic forms of statistical analysis. Thanks to improved analytical methods, accurate and extensive quantitative measurements can now be made of the mRNA, protein and metabolite from biological systems. Here, we report a large-scale analysis of the population abundance distributions of the transcriptomes, proteomes and metabolomes from varied biological systems. RESULTS: We compared the observed empirical distributions with a number of distributions: power law, lognormal, loglogistic, loggamma, right Pareto-lognormal (PLN) and double PLN (dPLN). The best-fit for mRNA, protein and metabolite population abundance distributions was found to be the dPLN. This distribution behaves like a lognormal distribution around the centre, and like a power law distribution in the tails. To better understand the cause of this observed distribution, we explored a simple stochastic model based on geometric Brownian motion. The distribution indicates that multiplicative effects are causally dominant in biological systems. We speculate that these effects arise from chemical reactions: the central-limit theorem then explains the central lognormal, and a number of possible mechanisms could explain the long tails: positive feedback, network topology, etc. Many of the components in the central lognormal parts of the empirical distributions are unidentified and/or have unknown function. This indicates that much more biology awaits discovery.
Chuan Lu, Ross D. King
Bioinform.2
2008 The EXACT description of biomedical protocols
abstract
MOTIVATION: Many published manuscripts contain experiment protocols which are poorly described or deficient in information. This means that the published results are very hard or impossible to repeat. This problem is being made worse by the increasing complexity of high-throughput/automated methods. There is therefore a growing need to represent experiment protocols in an efficient and unambiguous way. RESULTS: We have developed the Experiment ACTions (EXACT) ontology as the basis of a method of representing biological laboratory protocols. We provide example protocols that have been formalized using EXACT, and demonstrate the advantages and opportunities created by using this formalization. We argue that the use of EXACT will result in the publication of protocols with increased clarity and usefulness to the scientific community. AVAILABILITY: The ontology, examples and code can be downloaded from http://www.aber.ac.uk/compsci/Research/bio/dss/EXACT/.
Larisa N. Soldatova, Wayne Aubrey, Ross D. King, Amanda Clare
ISMB3
2008 Using a logical model to predict the growth of yeast
abstract
BACKGROUND: A logical model of the known metabolic processes in S. cerevisiae was constructed from iFF708, an existing Flux Balance Analysis (FBA) model, and augmented with information from the KEGG online pathway database. The use of predicate logic as the knowledge representation for modelling enables an explicit representation of the structure of the metabolic network, and enables logical inference techniques to be used for model identification/improvement. RESULTS: Compared to the FBA model, the logical model has information on an additional 263 putative genes and 247 additional reactions. The correctness of this model was evaluated by comparison with iND750 (an updated FBA model closely related to iFF708) by evaluating the performance of both models on predicting empirical minimal medium growth data/essential gene listings. CONCLUSION: ROC analysis and other statistical studies revealed that use of the simpler logical form and larger coverage results in no significant degradation of performance compared to iND750.
Ken E. Whelan, Ross D. King
BMC Bioinform.2
2008 Qualitative System Identification from Imperfect Data
abstract
Experience in the physical sciences suggests that the only realistic means of understanding complex systems is through the use of mathematical models. Typically, this has come to mean the identification of quantitative models expressed as differential equations. Quantitative modelling works best when the structure of the model (i.e., the form of the equations) is known; and the primary concern is one of estimating the values of the parameters in the model. For complex biological systems, the model-structure is rarely known and the modeler has to deal with both model-identification and parameter-estimation. In this paper we are concerned with providing automated assistance to the first of these problems. Specifically, we examine the identification by machine of the structural relationships between experimentally observed variables. These relationship will be expressed in the form of qualitative abstractions of a quantitative model. Such qualitative models may not only provide clues to the precise quantitative model, but also assist in understanding the essence of that model. Our position in this paper is that background knowledge incorporating system modelling principles can be used to constrain effectively the set of good qualitative models. Utilising the model-identification framework provided by Inductive Logic Programming (ILP) we present empirical support for this position using a series of increasingly complex artificial datasets. The results are obtained with qualitative and quantitative data subject to varying amounts of noise and different degrees of sparsity. The results also point to the presence of a set of qualitative states, which we term kernel subsets, that may be necessary for a qualitative model-learner to learn correct models. We demonstrate scalability of the method to biological system modelling by identification of the glycolysis metabolic pathway from data.
George Macleod Coghill, Ashwin Srinivasan 0001, Ross D. King
J. Artif. Intell. Res.3
2008 Incremental Identification of Qualitative Models of Biological Systems using Inductive Logic Programming
Ashwin Srinivasan 0001, Ross D. King
J. Mach. Learn. Res.2
2007 Evolutionary Optimization of Three-Photon Absorption in Molecular Iodine
abstract
We report on the application of an evolutionary algorithm to a noisy, dynamic optimization problem in chemistry: the maximization of three-photon absorption in molecular iodine. An evolution strategy is used in real-time in a closed loop experiment to search the space of physically realizable phase-modulated femtosecond laser pulses. The probability of three-photon absorption is estimated by measuring UV fluorescence. With the evolutionary search it is possible to enhance the UV fluorescence by a factor of 3.4 compared to the most intense pulse.
Robert Burbidge, Jem J. Rowland, Ross D. King, Nicholas T. Form, Benjamin J. Whitaker
CIDM3
2007 Active Learning for Regression Based on Query by Committee
Robert Burbidge, Jem J. Rowland, Ross D. King
IDEAL3
2007 Locational distribution of gene functional classes in Arabidopsis thaliana
abstract
BACKGROUND: We are interested in understanding the locational distribution of genes and their functions in genomes, as this distribution has both functional and evolutionary significance. Gene locational distribution is known to be affected by various evolutionary processes, with tandem duplication thought to be the main process producing clustering of homologous sequences. Recent research has found clustering of protein structural families in the human genome, even when genes identified as tandem duplicates have been removed from the data. However, this previous research was hindered as they were unable to analyse small sample sizes. This is a challenge for bioinformatics as more specific functional classes have fewer examples and conventional statistical analyses of these small data sets often produces unsatisfactory results. RESULTS: We have developed a novel bioinformatics method based on Monte Carlo methods and Greenwood's spacing statistic for the computational analysis of the distribution of individual functional classes of genes (from GO). We used this to make the first comprehensive statistical analysis of the relationship between gene functional class and location on a genome. Analysis of the distribution of all genes except tandem duplicates on the five chromosomes of A. thaliana reveals that the distribution on chromosomes I, II, IV and V is clustered at P = 0.001. Many functional classes are clustered, with the degree of clustering within an individual class generally consistent across all five chromosomes. A novel and surprising result was that the locational distribution of some functional classes were significantly more evenly spaced than would be expected by chance. CONCLUSION: Analysis of the A. thaliana genome reveals evidence of unexplained order in the locational distribution of genes. The same general analysis method can be applied to any genome, and indeed any sequential data involving classes.
Michael C. Riley, Amanda Clare, Ross D. King
BMC Bioinform.3
2006 Functional bioinformatics for Arabidopsis thaliana
abstract
MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.
Amanda Clare, Andreas Karwath, Helen Ougham, Ross D. King
Bioinform.4
2006 Functional bioinformatics for Arabidopsis thaliana
abstract
The authors would like to apologise for an error in citing H. Ougham's affiliated institution. It should be IGER, Plas Gogerddan, Aberystwyth, Ceredigion SY23 3EB, UK.
Amanda Clare, Andreas Karwath, Helen Ougham, Ross D. King
Bioinform.4
2006 Guest editorial
Rui Camacho, Ross D. King, Ashwin Srinivasan 0001
Mach. Learn.2
2006 Quantitative pharmacophore models with inductive logic programming
Ashwin Srinivasan 0001, David Page, Rui Camacho, Ross D. King
Mach. Learn.4
2005 The Robot Scientist Project
Ross D. King
ALT1
2005 The Robot Scientist Project
Ross D. King, Amanda Clare, Kenneth Whelan, Jem J. Rowland
Discovery Science1
2005 On the use of qualitative reasoning to simulate and identify metabolic pathway
abstract
MOTIVATION: Perhaps the greatest challenge of modern biology is to develop accurate in silico models of cells. To do this we require computational formalisms for both simulation (how according to the model the state of the cell evolves over time) and identification (learning a model cell from observation of states). We propose the use of qualitative reasoning (QR) as a unified formalism for both tasks. The two most commonly used alternative methods of modelling biochemical pathways are ordinary differential equations (ODEs), and logical/graph-based (LG) models. RESULTS: The QR formalism we use is an abstraction of ODEs. It enables the behaviour of many ODEs, with different functional forms and parameters, to be captured in a single QR model. QR has the advantage over LG models of explicitly including dynamics. To simulate biochemical pathways we have developed 'enzyme' and 'metabolite' QR building blocks that fit together to form models. These models are finite, directly executable, easy to interpret and robust. To identify QR models we have developed heuristic chemoinformatics graph analysis and machine learning procedures. The graph analysis procedure is a series of constraints and heuristics that limit the number of ways metabolites can combine to form pathways. The machine learning procedure is generate-and-test inductive logic programming. We illustrate the use of QR for modelling and simulation using the example of glycolysis. AVAILABILITY: All data and programs used are available on request.
Ross D. King, Simon M. Garrett, George Macleod Coghill
Bioinform.1
2005 A Dichotomic Search Algorithm for Mining and Learning in Domain-Specific Logics
Sébastien Ferré, Ross D. King
Fundam. Informaticae2
2004 Learning Qualitative Metabolic Models
George Macleod Coghill, Simon M. Garrett, Ross D. King
ECAI3
2004 BLID: An Application of Logical Information Systems to Bioinformatics
Sébastien Ferré, Ross D. King
ICFCA2
2004 Poly-transformation
Ross D. King, Mohammed Ouali
IDEAL1
2004 Confirmation of data mining based predictions of protein function
abstract
MOTIVATION: A central problem in bioinformatics is the assignment of function to sequenced open reading frames (ORFs). The most common approach is based on inferred homology using a statistically based sequence similarity (SIM) method, e.g. PSI-BLAST. Alternative non-SIM based bioinformatic methods are becoming popular. One such method is Data Mining Prediction (DMP). This is based on combining evidence from amino-acid attributes, predicted structure and phylogenic patterns; and uses a combination of Inductive Logic Programming data mining, and decision trees to produce prediction rules for functional class. DMP predictions are more general than is possible using homology. In 2000/1, DMP was used to make public predictions of the function of 1309 Escherichia coli ORFs. Since then biological knowledge has advanced allowing us to test our predictions. RESULTS: We examined the updated (20.02.02) Riley group genome annotation, and examined the scientific literature for direct experimental derivations of ORF function. Both tests confirmed the DMP predictions. Accuracy varied between rules, and with the detail of prediction, but they were generally significantly better than random. For voting rules, accuracies of 75-100% were obtained. Twenty-one of these DMP predictions have been confirmed by direct experimentation. The DMP rules also have interesting biological explanations. DMP is, to the best of our knowledge, the first non-SIM based prediction method to have been tested directly on new data. AVAILABILITY: We have designed the "Genepredictions" database for protein functional predictions. This is intended to act as an open repository for predictions for any organism and can be accessed at http://www.genepredictions.org
Ross D. King, Paul H. Wise, Amanda Clare
Bioinform.1
2003 A Personal View of How Best to Apply ILP
Ross D. King
ILP1
2003 Data Mining the Yeast Genome in a Lazy Functional Language
Amanda Clare, Ross D. King
PADL2
2003 Application of Inductive Logic Programming to Structure-Based Drug Design
David P. Enot, Ross D. King
PKDD2
2003 Statistical Evaluation of the Predictive Toxicology Challenge 2000-2001
abstract
MOTIVATION: The development of in silico models to predict chemical carcinogenesis from molecular structure would help greatly to prevent environmentally caused cancers. The Predictive Toxicology Challenge (PTC) competition was organized to test the state-of-the-art in applying machine learning to form such predictive models. RESULTS: Fourteen machine learning groups generated 111 models. The use of Receiver Operating Characteristic (ROC) space allowed the models to be uniformly compared regardless of the error cost function. We developed a statistical method to test if a model performs significantly better than random in ROC space. Using this test as criteria five models performed better than random guessing at a significance level p of 0.05 (not corrected for multiple testing). Statistically the best predictor was the Viniti model for female mice, with p value below 0.002. The toxicologically most interesting models were Leuven2 for male mice, and Kwansei for female rats. These models performed well in the statistical analysis and they are in the middle of ROC space, i.e. distant from extreme cost assumptions. These predictive models were also independently judged by domain experts to be among the three most interesting, and are believed to include a small but significant amount of empirically learned toxicological knowledge. AVAILABILITY: PTC details and data can be found at: http://www.predictive-toxicology.org/ptc/.
Hannu Toivonen, Ashwin Srinivasan 0001, Ross D. King, Stefan Kramer 0001, Christoph Helma
Bioinform.3
2003 An Empirical Study of the Use of Relevance Information in Inductive Logic Programming
Ashwin Srinivasan 0001, Ross D. King, Michael Bain 0001
J. Mach. Learn. Res.2
2002 Machine learning of functional class from phenotype data
abstract
MOTIVATION: Mutant phenotype growth experiments are an important novel source of functional genomics data which have received little attention in bioinformatics. We applied supervised machine learning to the problem of using phenotype data to predict the functional class of Open Reading Frames (ORFs) in Saccaromyces cerevisiae. Three sources of data were used: TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces (TRIPLES), European Functional Analysis Network (EUROFAN) and Munich Information Center for Protein Sequences (MIPS). The analysis of the data presented a number of challenges to machine learning: multi-class labels, a large number of sparsely populated classes, the need to learn a set of accurate rules (not a complete classification), and a very large amount of missing values. We modified the algorithm C4.5 to deal with these problems. RESULTS: Rules were learnt which are accurate and biologically meaningful. The rules predict function of 83 ORFs of unknown function at an estimated accuracy of > or = 80%.
Amanda Clare, Ross D. King
Bioinform.2
2002 Homology Induction: the use of machine learning to improve sequence similarity searches
abstract
BACKGROUND: The inference of homology between proteins is a key problem in molecular biology The current best approaches only identify approximately 50% of homologies (with a false positive rate set at 1/1000). RESULTS: We present Homology Induction (HI), a new approach to inferring homology. HI uses machine learning to bootstrap from standard sequence similarity search methods. First a standard method is run, then HI learns rules which are true for sequences of high similarity to the target (assumed homologues) and not true for general sequences, these rules are then used to discriminate sequences in the twilight zone. To learn the rules HI describes the sequences in a novel way based on a bioinformatic knowledge base, and the machine learning method of inductive logic programming. To evaluate HI we used the PDB40D benchmark which lists sequences of known homology but low sequence similarity. We compared the HI methodology with PSI-BLAST alone and found HI performed significantly better. In addition, Receiver Operating Characteristic (ROC) curve analysis showed that these improvements were robust for all reasonable error costs. The predictive homology rules learnt by HI by can be interpreted biologically to provide insight into conserved features of homologous protein families. CONCLUSIONS: HI is a new technique for the detection of remote protein homology--a central bioinformatic problem. HI with PSI-BLAST is shown to outperform PSI-BLAST for all error costs. It is expect that similar improvements would be obtained using HI with any sequence similarity method.
Andreas Karwath, Ross D. King
BMC Bioinform.2
2001 An Automated ILP Server in the Field of Bioinformatics
Andreas Karwath, Ross D. King
ILP2
2001 Knowledge Discovery in Multi-label Phenotype Data
Amanda Clare, Ross D. King
PKDD2
2001 The Predictive Toxicology Challenge 2000-2001
abstract
Abstract Summary: We initiated the Predictive Toxicology Challenge (PTC) to stimulate the development of advanced SAR techniques for predictive toxicology models. The goal of this challenge is to predict the rodent carcinogenicity of new compounds based on the experimental results of the US National Toxicology Program (NTP). Submissions will be evaluated on quantitative and qualitative scales to select the most predictive models and those with the highest toxicological relevance. Availability: http://www.informatik.uni-freiburg.de/~ml/ptc/ Contact: [email protected] * To whom correspondence should be addressed.
Christoph Helma, Ross D. King, Stefan Kramer 0001, Ashwin Srinivasan 0001
Bioinform.2
2001 The utility of different representations of protein sequence for predicting functional class
abstract
Abstract Motivation: Data Mining Prediction (DMP) is a novel approach to predicting protein functional class from sequence. DMP works even in the absence of a homologous protein of known function. We investigate the utility of different ways of representing protein sequence in DMP (residue frequencies, phylogeny, predicted structure) using the Escherichia coli genome as a model. Results: Using the different representations DMP learnt prediction rules that were more accurate than default at every level of function using every type of representation. The most effective way to represent sequence was using phylogeny (75% accuracy and 13% coverage of unassigned ORFs at the most general level of function: 69% accuracy and 7% coverage at the most detailed). We tested different methods for combining predictions from the different types of representation. These improved both the accuracy and coverage of predictions, e.g. 40% of all unassigned ORFs could be predicted at an estimated accuracy of 60% and 5% of unassigned ORFs could be predicted at an estimated accuracy of 86%. Availability: The rules and data are freely available. Warmr is free to academics. Contact: [email protected] Supplementary information: http://www.aber.ac.uk/~dcswww/Research/bio/ProteinFunction * To whom correspondence should be addressed.
Ross D. King, Andreas Karwath, Amanda Clare, Luc Dehaspe
Bioinform.1
2000 Genome scale prediction of protein functional class from sequence using data mining
abstract
The ability to predict protein function from amino acid sequence is a central research goal of molecular biology.Such a capability would greatly aid the biological interpretation of the genomic data and accelerate its medical exploitation.For the existing sequenced genomes function can be assigned to typically only between 40-60% of the genes [4,8,12,7].The new science of functional genomics is dedicated to discovering the function of these genes, and to further detailing gene function [10,27,17,6].Here we present a novel data-mining [24,18] approach to predicting protein functional class from sequence.We demonstrate the effectiveness of this approach on the Mycobacterium tuberculosis [8] genome.Biologically interpretable rules are identified that can predict protein function even in the absence of identifiable sequence homology.These rules predict 65% of the genes with no previous assigned function in Mycobacterium tuberculosis (the bacteria which causes TB) with an estimated accuracy of 60-80% (depending on the level of functional assignment).The rules give insight into the evolutionary history of the organism.
Ross D. King, Andreas Karwath, Amanda Clare, Luc Dehaspe
KDD1
1999 An assessment of submissions made to the Predictive Toxicology Evaluation Challenge
Ashwin Srinivasan 0001, Ross D. King, Douglas W. Bristol
IJCAI2
1999 Feature Construction with Inductive Logic Programming: A Study of Quantitative Predictions of Biological Activity Aided by Structural Attributes
Ashwin Srinivasan 0001, Ross D. King
Data Min. Knowl. Discov.2
1998 Biochemical Knowledge Discovery Using Inductive Logic Programming
Stephen H. Muggleton, Ashwin Srinivasan 0001, Ross D. King, Michael J. E. Sternberg
Discovery Science3
1998 Finding Frequent Substructures in Chemical Compounds
Luc Dehaspe, Hannu Toivonen, Ross D. King
KDD3
1997 The Predictive Toxicology Evaluation Challenge
Ashwin Srinivasan 0001, Ross D. King, Stephen H. Muggleton, Michael J. E. Sternberg
IJCAI (1)2
1997 DSC: public domain protein secondary structure predication
abstract
Ross D. King, Mansoor Saqi, Roger Sayle, Michael J.E. Sternberg; DSC: public domain protein secondary structure prediction, Bioinformatics, Volume 13, Issue 4,
Ross D. King, Mansoor A. S. Saqi, Roger A. Sayle, Michael J. E. Sternberg
Comput. Appl. Biosci.1
1996 Theories for Mutagenicity: A Study in First-Order and Feature-Based Induction
Ashwin Srinivasan 0001, Stephen H. Muggleton, Michael J. E. Sternberg, Ross D. King
Artif. Intell.4
1996 PM - protein music
abstract
We present the program PM for the analysis of protein sequence information using audification. Audification is the technique of using the sense of hearing to analyse data (Jackson, 1993). The advantage of audification over visualisation in data analysis is that sound has the property that when different notes are played together they can still be individually heard: in vision colours blend to form new colours. This distinction can be very useful when studying multivariate data. The one dimensional structure of DNA maps naturally onto musical sequences (Hayashi and Munakata, 1984; Ohno and Ohno, 1986). The mapping algorithm in PM is designed to translate coding sequences into sound in a way that is more harmonious and more open to analysis than previous methods. The PM algorithm uses the DNA nucleotide sequence to form the musical top line, and the properties of the translated amino-acids to form the bass line. The notes are in the scale of C Major and there are three beats to the bar. Each codon corresponds to a bar of music and the notes are played on the beat in order of nucleotide sequence. The mapping of the top line to musical notes is as follows: A4 submediant E3 mediant C3 tonic G3 dominant The mapping of the bass line to notes is more complicated and is based on the designation of aminoacid physico-chemical properties of Taylor (1983). Each of the properties has a note assigned to it. The mapping is as follows: polar hydrophobic charged positive aliphatic aromatic tiny
Ross D. King, C. G. Angus
Comput. Appl. Biosci.1
1994 Inductive Logic Programming Used to Discover Topological Constraints in Protein Structures
Ross D. King, Dominic A. Clark, Jack Shirazi, Michael J. E. Sternberg
ISMB1
1992 Statistical Methods in Learning
A. Sutherland, Bob Henery, Rafael Molina 0001, Charles C. Taylor, Ross D. King
IPMU5