EDBT 2026 Demo / reviewers in the wild / expert
Benjamin Goudey
dblp:123/7243 · also Benjamin W. Goudey
· DBLP profile ↗
11ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-2318-985XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A survey on deep learning for polygenic risk scoresabstractPolygenic risk scores (PRS) combine the effects of multiple genetic variants to predict an individual's genetic predisposition to a disease. PRS typically rely on linear models, which assume that all genetic variants act independently. They often fall short in predictive accuracy and are not able to explain the genetic variability of a trait to the full extent. There is growing interest in applying deep learning neural networks to model PRS given their ability to model non-linear relationships and strong performance in other domains. We conducted a survey of the literature to investigate how neural networks model PRS. We categorize deep learning-based approaches by their underlying architecture, highlighting their modeling assumptions, likely strengths and potential weaknesses of the architectures. Several categories of neural network architectures exhibited promising signs for the improvement of PRS' predictive power, namely sequence-based architectures, graph neural networks and those that incorporated biological knowledge. Additionally, the use of latent representations in autoencoders has improved predictive performance across diverse ancestries. However, a lack of existing model benchmarks on consistent datasets and phenotypes makes it challenging to understand the extent to which different architectures improve performance. Interpretability of deep learning-based PRS is also challenging with great care required when inferring causation. To address these challenges, we suggest the establishment and adherence to reporting standards and benchmarks to aid the development of deep learning-based PRS to find quantifiable trends in neural network architectures. Max Schuran, Benjamin Goudey, Gillian S. Dite, Enes Makalic |
Briefings Bioinform. | 2 |
| 2024 | Integration of background knowledge for automatic detection of inconsistencies in gene ontology annotationabstractMOTIVATION: Biological background knowledge plays an important role in the manual quality assurance (QA) of biological database records. One such QA task is the detection of inconsistencies in literature-based Gene Ontology Annotation (GOA). This manual verification ensures the accuracy of the GO annotations based on a comprehensive review of the literature used as evidence, Gene Ontology (GO) terms, and annotated genes in GOA records. While automatic approaches for the detection of semantic inconsistencies in GOA have been developed, they operate within predetermined contexts, lacking the ability to leverage broader evidence, especially relevant domain-specific background knowledge. This paper investigates various types of background knowledge that could improve the detection of prevalent inconsistencies in GOA. In addition, the paper proposes several approaches to integrate background knowledge into the automatic GOA inconsistency detection process. RESULTS: We have extended a previously developed GOA inconsistency dataset with several kinds of GOA-related background knowledge, including GeneRIF statements, biological concepts mentioned within evidence texts, GO hierarchy and existing GO annotations of the specific gene. We have proposed several effective approaches to integrate background knowledge as part of the automatic GOA inconsistency detection process. The proposed approaches can improve automatic detection of self-consistency and several of the most prevalent types of inconsistencies. This is the first study to explore the advantages of utilizing background knowledge and to propose a practical approach to incorporate knowledge in automatic GOA inconsistency detection. We establish a new benchmark for performance on this task. Our methods may be applicable to various tasks that involve incorporating biological background knowledge. AVAILABILITY AND IMPLEMENTATION: https://github.com/jiyuc/de-inconsistency. Jiyu Chen, Benjamin Goudey, Nicholas Geard, Karin Verspoor |
Bioinform. | 2 |
| 2022 | Propagation, detection and correction of errors using the sequence database networkabstractNucleotide and protein sequences stored in public databases are the cornerstone of many bioinformatics analyses. The records containing these sequences are prone to a wide range of errors, including incorrect functional annotation, sequence contamination and taxonomic misclassification. One source of information that can help to detect errors are the strong interdependency between records. Novel sequences in one database draw their annotations from existing records, may generate new records in multiple other locations and will have varying degrees of similarity with existing records across a range of attributes. A network perspective of these relationships between sequence records, within and across databases, offers new opportunities to detect-or even correct-erroneous entries and more broadly to make inferences about record quality. Here, we describe this novel perspective of sequence database records as a rich network, which we call the sequence database network, and illustrate the opportunities this perspective offers for quantification of database quality and detection of spurious entries. We provide an overview of the relevant databases and describe how the interdependencies between sequence records across these databases can be exploited by network analyses. We review the process of sequence annotation and provide a classification of sources of error, highlighting propagation as a major source. We illustrate the value of a network perspective through three case studies that use network analysis to detect errors, and explore the quality and quantity of critical relationships that would inform such network analyses. This systematic description of a network perspective of sequence database records provides a novel direction to combat the proliferation of errors within these critical bioinformatics resources. Benjamin Goudey, Nicholas Geard, Karin Verspoor, Justin Zobel |
Briefings Bioinform. | 1 |
| 2022 | Exploring automatic inconsistency detection for literature-based gene ontology annotationabstractMOTIVATION: Literature-based gene ontology annotations (GOA) are biological database records that use controlled vocabulary to uniformly represent gene function information that is described in the primary literature. Assurance of the quality of GOA is crucial for supporting biological research. However, a range of different kinds of inconsistencies in between literature as evidence and annotated GO terms can be identified; these have not been systematically studied at record level. The existing manual-curation approach to GOA consistency assurance is inefficient and is unable to keep pace with the rate of updates to gene function knowledge. Automatic tools are therefore needed to assist with GOA consistency assurance. This article presents an exploration of different GOA inconsistencies and an early feasibility study of automatic inconsistency detection. RESULTS: We have created a reliable synthetic dataset to simulate four realistic types of GOA inconsistency in biological databases. Three automatic approaches are proposed. They provide reasonable performance on the task of distinguishing the four types of inconsistency and are directly applicable to detect inconsistencies in real-world GOA database records. Major challenges resulting from such inconsistencies in the context of several specific application settings are reported. This is the first study to introduce automatic approaches that are designed to address the challenges in current GOA quality assurance workflows. The data underlying this article are available in Github at https://github.com/jiyuc/AutoGOAConsistency. Jiyu Chen, Benjamin Goudey, Justin Zobel, Nicholas Geard, Karin Verspoor |
Bioinform. | 2 |
| 2021 | eQTLHap: a tool for comprehensive eQTL analysis considering haplotypic and genotypic effectsabstractMOTIVATION: The high accuracy of recent haplotype phasing tools is enabling the integration of haplotype (or phase) information more widely in genetic investigations. One such possibility is phase-aware expression quantitative trait loci (eQTL) analysis, where haplotype-based analysis has the potential to detect associations that may otherwise be missed by standard SNP-based approaches. RESULTS: We present eQTLHap, a novel method to investigate associations between gene expression and genetic variants, considering their haplotypic and genotypic effect. Using multiple simulations based on real data, we demonstrate that phase-aware eQTL analysis significantly outperforms typical SNP-based methods when the causal genetic architecture involves multiple SNPs. We show that phase-aware eQTL analysis is robust to phasing errors, showing only a minor impact ($<4\%$) on sensitivity. Applying eQTLHap to real GEUVADIS and GTEx datasets detects numerous novel eQTLs undetected by a single-SNP approach, with 22 eQTLs replicating across studies or tissue types, highlighting the utility of phase-aware eQTL analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/ziadbkh/eQTLHap. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online. Ziad Al Bkhetan, Gursharan Chana, Cheng Soon Ong, Benjamin Goudey, Kotagiri Ramamohanarao |
Briefings Bioinform. | 4 |
| 2021 | Evaluation of consensus strategies for haplotype phasingabstractHaplotype phasing is a critical step for many genetic applications but incorrect estimates of phase can negatively impact downstream analyses. One proposed strategy to improve phasing accuracy is to combine multiple independent phasing estimates to overcome the limitations of any individual estimate. However, such a strategy is yet to be thoroughly explored. This study provides a comprehensive evaluation of consensus strategies for haplotype phasing. We explore the performance of different consensus paradigms, and the effect of specific constituent tools, across several datasets with different characteristics and their impact on the downstream task of genotype imputation. Based on the outputs of existing phasing tools, we explore two different strategies to construct haplotype consensus estimators: voting across outputs from multiple phasing tools and multiple outputs of a single non-deterministic tool. We find that the consensus approach from multiple tools reduces SE by an average of 10% compared to any constituent tool when applied to European populations and has the highest accuracy regardless of population ethnicity, sample size, variant density or variant frequency. Furthermore, the consensus estimator improves the accuracy of the downstream task of genotype imputation carried out by the widely used Minimac3, pbwt and BEAGLE5 tools. Our results provide guidance on how to produce the most accurate phasing estimates and the trade-offs that a consensus approach may have. Our implementation of consensus haplotype phasing, consHap, is available freely at https://github.com/ziadbkh/consHap. Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. Ziad Al Bkhetan, Gursharan Chana, Kotagiri Ramamohanarao, Karin Verspoor, Benjamin Goudey |
Briefings Bioinform. | 5 |
| 2019 | Exploring effective approaches for haplotype block phasingabstractBACKGROUND: Knowledge of phase, the specific allele sequence on each copy of homologous chromosomes, is increasingly recognized as critical for detecting certain classes of disease-associated mutations. One approach for detecting such mutations is through phased haplotype association analysis. While the accuracy of methods for phasing genotype data has been widely explored, there has been little attention given to phasing accuracy at haplotype block scale. Understanding the combined impact of the accuracy of phasing tool and the method used to determine haplotype blocks on the error rate within the determined blocks is essential to conduct accurate haplotype analyses. RESULTS: We present a systematic study exploring the relationship between seven widely used phasing methods and two common methods for determining haplotype blocks. The evaluation focuses on the number of haplotype blocks that are incorrectly phased. Insights from these results are used to develop a haplotype estimator based on a consensus of three tools. The consensus estimator achieved the most accurate phasing in all applied tests. Individually, EAGLE2, BEAGLE and SHAPEIT2 alternate in being the best performing tool in different scenarios. Determining haplotype blocks based on linkage disequilibrium leads to more correctly phased blocks compared to a sliding window approach. We find that there is little difference between phasing sections of a genome (e.g. a gene) compared to phasing entire chromosomes. Finally, we show that the location of phasing error varies when the tools are applied to the same data several times, a finding that could be important for downstream analyses. CONCLUSIONS: The choice of phasing and block determination algorithms and their interaction impacts the accuracy of phased haplotype blocks. This work provides guidance and evidence for the different design choices needed for analyses using haplotype blocks. The study highlights a number of issues that may have limited the replicability of previous haplotype analysis. Ziad Al Bkhetan, Justin Zobel, Adam Kowalczyk, Karin Verspoor, Benjamin Goudey |
BMC Bioinform. | 5 |
| 2018 | A hybrid approach for automated mutation annotation of the extended human mutation landscape in scientific literature
Antonio Jimeno-Yepes, Andrew MacKinlay, Natalie Gunn, Christine Schieber, Noel Faux, Matthew Downton, Benjamin Goudey |
AMIA | 7 |
| 2018 | Predicting the Key Alzheimers Biomarkers in CSF from Plasma Analytes
Christine Schieber, Annalisa Swan, Roslyn I. Hickson, Noel Faux, Benjamin Goudey |
AMIA | 5 |
| 2014 | GWISFI: A universal GPU interface for exhaustive search of pairwise interactions in case-control GWAS in minutesabstractEpistatic interactions between genes are believed to be a critical component in the genetic architecture of complex diseases. Genome Wide Association Studies (GWAS) may be able to detect such genetic interactions indirectly, via the identification of associated SNP markers. Major obstacles to progress in this area are: the unknown nature of epistatic interactions, little understanding of the capabilities of different filtering methods, and the computational difficulties for exhaustive analysis. A common platform enabling various detection methods is needed to avoid practical issues such as software compatibility and portability, incompatible input and output formats and varying demands on computational resources. We developed a highly optimised GPU system capable of exhaustively analysing all SNP-pairs in typical GWAS data (0.5M SNPs, 5K samples) in a few minutes on a standard desktop computer. A number of programming elements provided by a functional interface can be used to construct user-defined statistical tests to efficiently score every SNP pair. As a proof of principle, we have implemented 8 methods from the literature via our interface. We have applied all of them using a single GPU to exhaustively scan the 7 popular WTCCC case-control GWAS datasets. We present timing results for these methods, both in their original software implementations and using our platform. Significant improvements in timing are observed, up to 10000 times for CPU implementations of the popular FastEpistasis in PLINK and up to 2 orders of magnitude for some GPU implementations in the literature. As an initial discovery we show plots for overlaps of list of selected pairs by 8 algorithms for Type 2 Diabetes, WTCCC data. Andrew Kowalczyk, Richard M. Campbell, Benjamin Goudey, David Rawlinson 0001, Aaron Harwood, Herman L. Ferrá, Adam Kowalczyk |
BIBM | 5 |
| 2011 | Replication of epistatic DNA loci in two case-control GWAS studies using OPE algorithmabstractBackground One of the limiting factors of current genome-wide association studies (GWAS) is the inability of current methods to comprehensively examine SNP interactions for a reasonable sized dataset. It is hypothesised that this limitation is one of the reasons that GWAS studies have not been able to have a greater impact [1,2]. Many current methods for handling interactions are computationally expensive and do not scale to entire studies. Those methods that do scale often achieve this by pruning their datasets in some manner. This is commonly done by considering only those SNPs that show strong marginal effects, despite the fact that a strongly interacting pair may consist of SNPs with low effects individually. Benjamin Goudey, David Rawlinson 0001, Armita Zarnegar, Eder Kikianty, John Markham, Geoff MacIntyre, Gad Abraham, Linda Stern, Michael Inouye, Izhak Haviv, Adam Kowalczyk |
BMC Bioinform. | 1 |