Benjamin L. Elsworth

dblp:197/8223 · also Ben Elsworth · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-7328-4233ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Integrating Mendelian randomization and literature-mined evidence for breast cancer risk factors
abstract
OBJECTIVE: An increasing challenge in population health research is efficiently utilising the wealth of data available from multiple sources to investigate disease mechanisms and identify potential intervention targets. The use of biomedical data integration platforms can facilitate evidence triangulation from these different sources, improving confidence in causal relationships of interest. In this work, we aimed to integrate Mendelian randomization (MR) and literature-mined evidence from the EpiGraphDB biomedical knowledge graph to build a comprehensive overview of risk factors for developing breast cancer. METHODS: We utilised MR-EvE ("Everything-vs-Everything") data to identify candidate risk factors for breast cancer and generate hypotheses for potential mediators of their effect. We also integrated this data with literature-mined relationships, which were extracted by overlapping literature spaces of risk factors and breast cancer. The literature-based discovery (LBD) results were followed up by validation with two-step MR to triangulate the findings from two data sources. RESULTS: We identified 129 novel and established lifestyle risk factors and molecular traits with evidence of an effect on breast cancer, and made the MR results available in an R/Shiny app (https://mvab.shinyapps.io/MR_heatmaps/). We developed an LBD approach for identifying potential mechanistic intermediates of identified risk factors. We present the results of MR and literature evidence integration for two case studies (childhood body size and HDL-cholesterol), demonstrating their complementary functionalities. CONCLUSION: We demonstrate that MR-EvE data offers an efficient hypothesis-generating approach for identifying disease risk factors. Moreover, we show that integrating MR evidence with literature-mined data may be used to identify causal intermediates and uncover the mechanisms behind the disease.
Marina Vabistsevits, Timothy Robinson, Benjamin L. Elsworth, Yi Liu 0058, Tom R. Gaunt
J. Biomed. Informatics3
2023 Using language models and ontology topology to perform semantic mapping of traits between biomedical datasets
abstract
MOTIVATION: Human traits are typically represented in both the biomedical literature and large population studies as descriptive text strings. Whilst a number of ontologies exist, none of these perfectly represent the entire human phenome and exposome. Mapping trait names across large datasets is therefore time-consuming and challenging. Recent developments in language modelling have created new methods for semantic representation of words and phrases, and these methods offer new opportunities to map human trait names in the form of words and short phrases, both to ontologies and to each other. Here, we present a comparison between a range of established and more recent language modelling approaches for the task of mapping trait names from UK Biobank to the Experimental Factor Ontology (EFO), and also explore how they compare to each other in direct trait-to-trait mapping. RESULTS: In our analyses of 1191 traits from UK Biobank with manual EFO mappings, the BioSentVec model performed best at predicting these, matching 40.3% of the manual mappings correctly. The BlueBERT-EFO model (finetuned on EFO) performed nearly as well (38.8% of traits matching the manual mapping). In contrast, Levenshtein edit distance only mapped 22% of traits correctly. Pairwise mapping of traits to each other demonstrated that many of the models can accurately group similar traits based on their semantic similarity. AVAILABILITY AND IMPLEMENTATION: Our code is available at https://github.com/MRCIEU/vectology.
Yi Liu 0058, Benjamin L. Elsworth, Tom R. Gaunt
Bioinform.2
2021 MELODI Presto: a fast and agile tool to explore semantic triples derived from biomedical literature
abstract
SUMMARY: The field of literature-based discovery is growing in step with the volume of literature being produced. From modern natural language processing algorithms to high quality entity tagging, the methods and their impact are developing rapidly. One annotation object that arises from these approaches, the subject-predicate-object triple, is proving to be very useful in representing knowledge. We have implemented efficient search methods and an application programming interface, to create fast and convenient functions to utilize triples extracted from the biomedical literature by SemMedDB. By refining these data, we have identified a set of triples that focus on the mechanistic aspects of the literature, and provide simple methods to explore both enriched triples from single queries, and overlapping triples across two query lists. AVAILABILITY AND IMPLEMENTATION: https://melodi-presto.mrcieu.ac.uk/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Benjamin L. Elsworth, Tom R. Gaunt
Bioinform.1
2021 Erratum to: EpiGraphDB: a database and data mining platform for health data science
abstract
Bioinformatics (2020) doi: 10.1093/bioinformatics/btaa961 Upon the original publication of this article, there was an error in the source code syntax under sub-section “2.2 Integration of epidemiological evidence” in the “Materials and methods” section. The source code syntax should read: “(e.g. (Gwas {trait: ‘Body mass index’})-[MR {beta, se, pval}]->(Gwas {trait: ‘Coronary heart disease’}))” instead of “ (e.g. [Gwas (trait: ‘Body mass index’)]-[MR {beta, se, pval}]->(Gwas {trait: ‘Coronary heart disease’})))”. This error has now been corrected. The Publisher apologizes for the error.
Yi Liu 0058, Benjamin L. Elsworth, Pau Erola, Valeriia Haberland, Gibran Hemani, Matt Lyon, Jie Zheng 0006, Oliver Lloyd, Marina Vabistsevits, Tom R. Gaunt
Bioinform.2
2021 EpiGraphDB: a database and data mining platform for health data science
abstract
MOTIVATION: The wealth of data resources on human phenotypes, risk factors, molecular traits and therapeutic interventions presents new opportunities for population health sciences. These opportunities are paralleled by a growing need for data integration, curation and mining to increase research efficiency, reduce mis-inference and ensure reproducible research. RESULTS: We developed EpiGraphDB (https://epigraphdb.org/), a graph database containing an array of different biomedical and epidemiological relationships and an analytical platform to support their use in human population health data science. In addition, we present three case studies that illustrate the value of this platform. The first uses EpiGraphDB to evaluate potential pleiotropic relationships, addressing mis-inference in systematic causal analysis. In the second case study, we illustrate how protein-protein interaction data offer opportunities to identify new drug targets. The final case study integrates causal inference using Mendelian randomization with relationships mined from the biomedical literature to 'triangulate' evidence from different sources. AVAILABILITY AND IMPLEMENTATION: The EpiGraphDB platform is openly available at https://epigraphdb.org. Code for replicating case study results is available at https://github.com/MRCIEU/epigraphdb as Jupyter notebooks using the API, and https://mrcieu.github.io/epigraphdb-r using the R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yi Liu 0058, Benjamin L. Elsworth, Pau Erola, Valeriia Haberland, Gibran Hemani, Matt Lyon, Jie Zheng 0006, Oliver Lloyd, Marina Vabistsevits, Tom R. Gaunt
Bioinform.2
2017 LD Hub: a centralized database and web interface to perform LD score regression that maximizes the potential of summary level GWAS data for SNP heritability and genetic correlation analysis
abstract
MOTIVATION: LD score regression is a reliable and efficient method of using genome-wide association study (GWAS) summary-level results data to estimate the SNP heritability of complex traits and diseases, partition this heritability into functional categories, and estimate the genetic correlation between different phenotypes. Because the method relies on summary level results data, LD score regression is computationally tractable even for very large sample sizes. However, publicly available GWAS summary-level data are typically stored in different databases and have different formats, making it difficult to apply LD score regression to estimate genetic correlations across many different traits simultaneously. RESULTS: In this manuscript, we describe LD Hub - a centralized database of summary-level GWAS results for 173 diseases/traits from different publicly available resources/consortia and a web interface that automates the LD score regression analysis pipeline. To demonstrate functionality and validate our software, we replicated previously reported LD score regression analyses of 49 traits/diseases using LD Hub; and estimated SNP heritability and the genetic correlation across the different phenotypes. We also present new results obtained by uploading a recent atopic dermatitis GWAS meta-analysis to examine the genetic correlation between the condition and other potentially related traits. In response to the growing availability of publicly accessible GWAS summary-level results data, our database and the accompanying web interface will ensure maximal uptake of the LD score regression methodology, provide a useful database for the public dissemination of GWAS results, and provide a method for easily screening hundreds of traits for overlapping genetic aetiologies. AVAILABILITY AND IMPLEMENTATION: The web interface and instructions for using LD Hub are available at http://ldsc.broadinstitute.org/ CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Jie Zheng 0006, Mesut A. Erzurumluoglu, Benjamin L. Elsworth, John P. Kemp, Laurence J. Howe, Philip C. Haycock, Gibran Hemani, Katherine Tansey, Charles Laurin, Early Genetics, Beate St. Pourcain, Nicole M. Warrington, Hilary K. Finucane, Alkes L. Price, Brendan K. Bulik-Sullivan, Verneri Anttila, Lavinia Paternoster, Tom R. Gaunt, David M. Evans 0002, Benjamin M. Neale
Bioinform.3
2013 Badger - an accessible genome exploration environment
abstract
SUMMARY: High-quality draft genomes are now easy to generate, as sequencing and assembly costs have dropped dramatically. However, building a user-friendly searchable Web site and database for a newly annotated genome is not straightforward. Here we present Badger, a lightweight and easy-to-install genome exploration environment designed for next generation non-model organism genomes. AVAILABILITY: Badger is released under the GPL and is available at http://badger.bio.ed.ac.uk/. We show two working examples: (i) a test dataset included with the source code, and (ii) a collection of four filarial nematode genomes. CONTACT: [email protected].
Benjamin L. Elsworth, Martin O. Jones, Mark L. Blaxter
Bioinform.1