Ian P. Barrett

dblp:215/4590 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0005-7164-0376ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021
YearPublicationVenuePosition
2026 RP3Net: a deep learning model for predicting recombinant protein production in Escherichia coli
abstract
MOTIVATION: Recombinant protein expression can be a limiting step in the production of protein reagents for drug discovery and other biotechnology applications. We introduce RP3Net (Recombinant Protein Production Prediction Network), an AI model of small-scale heterologous soluble protein expression in Escherichia coli. RP3Net utilizes the most recent protein and genomic foundational models. A curated dataset of internal experimental results from AstraZeneca and publicly available data from the Structural Genomics Consortium was used for training, validation and testing of RP3Net. RESULTS: RP3Net achieves an increase in area under the receiver operator curve (AUROC) of 0.15, compared to a baseline model. When experimentally validated on an independent, prospective, manually selected set of 97 constructs, RP3Net outperformed currently available models, with an AUROC of 0.83, delivering accurate predictions in 77% of the cases, and correctly identifying successfully expressing constructs in 92% of cases. AVAILABILITY AND IMPLEMENTATION: The model, along with installation and running instructions, is available under an MIT licence at https://github.com/RP3Net/RP3Net, DOI 10.5281/zenodo.17243498.
Evgeny Tankhilevich, Sergio Martínez Cuesta, Ian P. Barrett, Carolina Berg, Lovisa Holmberg Schiavone, Andrew R. Leach
Bioinform.3
2025 The role of graph topology in the performance of biomedical Knowledge Graph Completion models
abstract
MOTIVATION: Knowledge Graph Completion has been increasingly adopted as a useful method for helping address several tasks in biomedical research, such as drug repurposing or drug-target identification. To that end, a variety of datasets and Knowledge Graph Embedding models have been proposed over the years. However, little is known about the properties that render a dataset, and associated modelling choices, useful for a given task. Moreover, even though theoretical properties of Knowledge Graph Embedding models are well understood, their practical utility in this field remains controversial. RESULTS: In this work, we conduct a comprehensive investigation into the topological properties of publicly available biomedical Knowledge Graphs and establish links to the accuracy observed in real-world tasks. By releasing all model predictions and a new suite of analysis tools we invite the community to build upon our work and continue improving the understanding of these crucial applications. AVAILABILITY AND IMPLEMENTATION: The code used to perform experiments and analyze results in this article as well as all experimental data is available at https://github.com/graphcore-research/kg-topology-toolbox/tree/main/the_role_of_graph_topology_paper and archived on Zenodo, at https://doi.org/10.5281/zenodo.12097376.
Alberto Cattaneo, Stephen Bonner, Thomas Martynec, Edward R. Morrissey, Carlo Luschi, Ian P. Barrett, Daniel Justus
Bioinform.6
2022 A review of biomedical datasets relating to drug discovery: a knowledge graph perspective
abstract
Drug discovery and development is a complex and costly process. Machine learning approaches are being investigated to help improve the effectiveness and speed of multiple stages of the drug discovery pipeline. Of these, those that use Knowledge Graphs (KG) have promise in many tasks, including drug repurposing, drug toxicity prediction and target gene-disease prioritization. In a drug discovery KG, crucial elements including genes, diseases and drugs are represented as entities, while relationships between them indicate an interaction. However, to construct high-quality KGs, suitable data are required. In this review, we detail publicly available sources suitable for use in constructing drug discovery focused KGs. We aim to help guide machine learning and KG practitioners who are interested in applying new techniques to the drug discovery field, but who may be unfamiliar with the relevant data sources. The datasets are selected via strict criteria, categorized according to the primary type of information contained within and are considered based upon what information could be extracted to build a KG. We then present a comparative analysis of existing public drug discovery KGs and an evaluation of selected motivating case studies from the literature. Additionally, we raise numerous and unique challenges and issues associated with the domain and its datasets, while also highlighting key future research directions. We hope this review will motivate KGs use in solving key and emerging questions in the drug discovery domain.
Stephen Bonner, Ian P. Barrett, Cheng Ye 0002, Rowan Swiers, Ola Engkvist, Andreas Bender 0002, Charles Tapley Hoyt, William L. Hamilton
Briefings Bioinform.2
2022 Implications of topological imbalance for representation learning on biomedical knowledge graphs
abstract
Adoption of recently developed methods from machine learning has given rise to creation of drug-discovery knowledge graphs (KG) that utilize the interconnected nature of the domain. Graph-based modelling of the data, combined with KG embedding (KGE) methods, are promising as they provide a more intuitive representation and are suitable for inference tasks such as predicting missing links. One common application is to produce ranked lists of genes for a given disease, where the rank is based on the perceived likelihood of association between the gene and the disease. It is thus critical that these predictions are not only pertinent but also biologically meaningful. However, KGs can be biased either directly due to the underlying data sources that are integrated or due to modeling choices in the construction of the graph, one consequence of which is that certain entities can get topologically overrepresented. We demonstrate the effect of these inherent structural imbalances, resulting in densely-connected entities being highly ranked no matter the context. We provide support for this observation across different datasets, models as well as predictive tasks. Further, we present various graph perturbation experiments which yield more support to the observation that KGE models can be more influenced by the frequency of entities rather than any biological information encoded within the relations. Our results highlight the importance of data modeling choices, and emphasizes the need for practitioners to be mindful of these issues when interpreting model outputs and during KG composition.
Stephen Bonner, Ufuk Kirik, Ola Engkvist, Ian P. Barrett
Briefings Bioinform.5
2022 A Knowledge Graph-Enhanced Tensor Factorisation Model for Discovering Drug Targets
abstract
The drug discovery and development process is a long and expensive one, costing over 1 billion USD on average per drug and taking 10-15 years. To reduce the high levels of attrition throughout the process, there has been a growing interest in applying machine learning methodologies to various stages of drug discovery and development in the recent decade, especially at the earliest stage - identification of druggable disease genes. In this paper, we have developed a new tensor factorisation model to predict potential drug targets (genes or proteins) for treating diseases. We created a three-dimensional data tensor consisting of 1,048 gene targets, 860 diseases and 230,011 evidence attributes and clinical outcomes connecting them, using data extracted from the Open Targets and PharmaProjects databases. We enriched the data with gene target representations learned from a drug discovery-oriented knowledge graph and applied our proposed method to predict the clinical outcomes for unseen gene target and disease pairs. We designed three evaluation strategies to measure the prediction performance and benchmarked several commonly used machine learning classifiers together with Bayesian matrix and tensor factorisation methods. The result shows that incorporating knowledge graph embeddings significantly improves the prediction accuracy and that training tensor factorisation alongside a dense neural network outperforms all other baselines. In summary, our framework combines two actively studied machine learning approaches to disease target identification, namely tensor factorisation and knowledge graph representation learning, which could be a promising avenue for further exploration in data-driven drug discovery.
Cheng Ye 0002, Rowan Swiers, Stephen Bonner, Ian P. Barrett
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 Orthologue chemical space and its influence on target prediction
abstract
Motivation: In silico approaches often fail to utilize bioactivity data available for orthologous targets due to insufficient evidence highlighting the benefit for such an approach. Deeper investigation into orthologue chemical space and its influence toward expanding compound and target coverage is necessary to improve the confidence in this practice. Results: Here we present analysis of the orthologue chemical space in ChEMBL and PubChem and its impact on target prediction. We highlight the number of conflicting bioactivities between human and orthologues is low and annotations are overall compatible. Chemical space analysis shows orthologues are chemically dissimilar to human with high intra-group similarity, suggesting they could effectively extend the chemical space modelled. Based on these observations, we show the benefit of orthologue inclusion in terms of novel target coverage. We also benchmarked predictive models using a time-series split and also using bioactivities from Chemistry Connect and HTS data available at AstraZeneca, showing that orthologue bioactivity inclusion statistically improved performance. Availability and implementation: Orthologue-based bioactivity prediction and the compound training set are available at www.github.com/lhm30/PIDGINv2. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Lewis H. Mervin, Krishna C. Bulusu, Leen Kalash, Avid M. Afzal, Fredrik Svensson, Mike A. Firth, Ian P. Barrett, Ola Engkvist, Andreas Bender 0002
Bioinform.7
2016 Kinase Inhibition Leads to Hormesis in a Dual Phosphorylation-Dephosphorylation Cycle
abstract
Many antimicrobial and anti-tumour drugs elicit hormetic responses characterised by low-dose stimulation and high-dose inhibition. While this can have profound consequences for human health, with low drug concentrations actually stimulating pathogen or tumour growth, the mechanistic understanding behind such responses is still lacking. We propose a novel, simple but general mechanism that could give rise to hormesis in systems where an inhibitor acts on an enzyme. At its core is one of the basic building blocks in intracellular signalling, the dual phosphorylation-dephosphorylation motif, found in diverse regulatory processes including control of cell proliferation and programmed cell death. Our analytically-derived conditions for observing hormesis provide clues as to why this mechanism has not been previously identified. Current mathematical models regularly make simplifying assumptions that lack empirical support but inadvertently preclude the observation of hormesis. In addition, due to the inherent population heterogeneities, the presence of hormesis is likely to be masked in empirical population-level studies. Therefore, examining hormetic responses at single-cell level coupled with improved mathematical models could substantially enhance detection and mechanistic understanding of hormesis.
Peter Rashkov, Ian P. Barrett, Robert E. Beardmore, Claus Bendtsen, Ivana Gudelj
PLoS Comput. Biol.2
2011 Automatic extraction of angiogenesis bioprocess from text
abstract
MOTIVATION: Understanding key biological processes (bioprocesses) and their relationships with constituent biological entities and pharmaceutical agents is crucial for drug design and discovery. One way to harvest such information is searching the literature. However, bioprocesses are difficult to capture because they may occur in text in a variety of textual expressions. Moreover, a bioprocess is often composed of a series of bioevents, where a bioevent denotes changes to one or a group of cells involved in the bioprocess. Such bioevents are often used to refer to bioprocesses in text, which current techniques, relying solely on specialized lexicons, struggle to find. RESULTS: This article presents a range of methods for finding bioprocess terms and events. To facilitate the study, we built a gold standard corpus in which terms and events related to angiogenesis, a key biological process of the growth of new blood vessels, were annotated. Statistics of the annotated corpus revealed that over 36% of the text expressions that referred to angiogenesis appeared as events. The proposed methods respectively employed domain-specific vocabularies, a manually annotated corpus and unstructured domain-specific documents. Evaluation results showed that, while a supervised machine-learning model yielded the best precision, recall and F1 scores, the other methods achieved reasonable performance and less cost to develop. AVAILABILITY: The angiogenesis vocabularies, gold standard corpus, annotation guidelines and software described in this article are available at http://text0.mib.man.ac.uk/~mbassxw2/angiogenesis/ CONTACT: [email protected].
Xinglong Wang, Iain McKendrick, Ian P. Barrett, Ian Dix, Tim French 0003, Jun'ichi Tsujii, Sophia Ananiadou
Bioinform.3