EDBT 2026 Demo / reviewers in the wild / expert
Nicholas Geard
dblp:68/6353
· DBLP profile ↗
20ranked-venue papers
8as first author
8since 2021 · last 2024
0000-0003-0069-2281ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Integration of background knowledge for automatic detection of inconsistencies in gene ontology annotationabstractMOTIVATION: Biological background knowledge plays an important role in the manual quality assurance (QA) of biological database records. One such QA task is the detection of inconsistencies in literature-based Gene Ontology Annotation (GOA). This manual verification ensures the accuracy of the GO annotations based on a comprehensive review of the literature used as evidence, Gene Ontology (GO) terms, and annotated genes in GOA records. While automatic approaches for the detection of semantic inconsistencies in GOA have been developed, they operate within predetermined contexts, lacking the ability to leverage broader evidence, especially relevant domain-specific background knowledge. This paper investigates various types of background knowledge that could improve the detection of prevalent inconsistencies in GOA. In addition, the paper proposes several approaches to integrate background knowledge into the automatic GOA inconsistency detection process. RESULTS: We have extended a previously developed GOA inconsistency dataset with several kinds of GOA-related background knowledge, including GeneRIF statements, biological concepts mentioned within evidence texts, GO hierarchy and existing GO annotations of the specific gene. We have proposed several effective approaches to integrate background knowledge as part of the automatic GOA inconsistency detection process. The proposed approaches can improve automatic detection of self-consistency and several of the most prevalent types of inconsistencies. This is the first study to explore the advantages of utilizing background knowledge and to propose a practical approach to incorporate knowledge in automatic GOA inconsistency detection. We establish a new benchmark for performance on this task. Our methods may be applicable to various tasks that involve incorporating biological background knowledge. AVAILABILITY AND IMPLEMENTATION: https://github.com/jiyuc/de-inconsistency. Jiyu Chen, Benjamin Goudey, Nicholas Geard, Karin Verspoor |
Bioinform. | 3 |
| 2024 | Ethical frameworks should be applied to computational modelling of infectious disease interventionsabstractThis perspective is part of an international effort to improve epidemiological models with the goal of reducing the unintended consequences of infectious disease interventions. The scenarios in which models are applied often involve difficult trade-offs that are well recognised in public health ethics. Unless these trade-offs are explicitly accounted for, models risk overlooking contested ethical choices and values, leading to an increased risk of unintended consequences. We argue that such risks could be reduced if modellers were more aware of ethical frameworks and had the capacity to explicitly account for the relevant values in their models. We propose that public health ethics can provide a conceptual foundation for developing this capacity. After reviewing relevant concepts in public health and clinical ethics, we discuss examples from the COVID-19 pandemic to illustrate the current separation between public health ethics and infectious disease modelling. We conclude by describing practical steps to build the capacity for ethically aware modelling. Developing this capacity constitutes a critical step towards ethical practice in computational modelling of public health interventions, which will require collaboration with experts on public health ethics, decision support, behavioural interventions, and social determinants of health, as well as direct consultation with communities and policy makers. Cameron Zachreson, Julian Savulescu, Freya M. Shearer, Michael J. Plank, Simon Coghlan, Joel C. Miller, Kylie E. C. Ainslie, Nicholas Geard |
PLoS Comput. Biol. | 8 |
| 2022 | Propagation, detection and correction of errors using the sequence database networkabstractNucleotide and protein sequences stored in public databases are the cornerstone of many bioinformatics analyses. The records containing these sequences are prone to a wide range of errors, including incorrect functional annotation, sequence contamination and taxonomic misclassification. One source of information that can help to detect errors are the strong interdependency between records. Novel sequences in one database draw their annotations from existing records, may generate new records in multiple other locations and will have varying degrees of similarity with existing records across a range of attributes. A network perspective of these relationships between sequence records, within and across databases, offers new opportunities to detect-or even correct-erroneous entries and more broadly to make inferences about record quality. Here, we describe this novel perspective of sequence database records as a rich network, which we call the sequence database network, and illustrate the opportunities this perspective offers for quantification of database quality and detection of spurious entries. We provide an overview of the relevant databases and describe how the interdependencies between sequence records across these databases can be exploited by network analyses. We review the process of sequence annotation and provide a classification of sources of error, highlighting propagation as a major source. We illustrate the value of a network perspective through three case studies that use network analysis to detect errors, and explore the quality and quantity of critical relationships that would inform such network analyses. This systematic description of a network perspective of sequence database records provides a novel direction to combat the proliferation of errors within these critical bioinformatics resources. Benjamin Goudey, Nicholas Geard, Karin Verspoor, Justin Zobel |
Briefings Bioinform. | 2 |
| 2022 | Exploring automatic inconsistency detection for literature-based gene ontology annotationabstractMOTIVATION: Literature-based gene ontology annotations (GOA) are biological database records that use controlled vocabulary to uniformly represent gene function information that is described in the primary literature. Assurance of the quality of GOA is crucial for supporting biological research. However, a range of different kinds of inconsistencies in between literature as evidence and annotated GO terms can be identified; these have not been systematically studied at record level. The existing manual-curation approach to GOA consistency assurance is inefficient and is unable to keep pace with the rate of updates to gene function knowledge. Automatic tools are therefore needed to assist with GOA consistency assurance. This article presents an exploration of different GOA inconsistencies and an early feasibility study of automatic inconsistency detection. RESULTS: We have created a reliable synthetic dataset to simulate four realistic types of GOA inconsistency in biological databases. Three automatic approaches are proposed. They provide reasonable performance on the task of distinguishing the four types of inconsistency and are directly applicable to detect inconsistencies in real-world GOA database records. Major challenges resulting from such inconsistencies in the context of several specific application settings are reported. This is the first study to introduce automatic approaches that are designed to address the challenges in current GOA quality assurance workflows. The data underlying this article are available in Github at https://github.com/jiyuc/AutoGOAConsistency. Jiyu Chen, Benjamin Goudey, Justin Zobel, Nicholas Geard, Karin Verspoor |
Bioinform. | 4 |
| 2022 | MPVNN: Mutated Pathway Visible Neural Network architecture for interpretable prediction of cancer-specific survival riskabstractMOTIVATION: Survival risk prediction using gene expression data is important in making treatment decisions in cancer. Standard neural network (NN) survival analysis models are black boxes with a lack of interpretability. More interpretable visible neural network architectures are designed using biological pathway knowledge. But they do not model how pathway structures can change for particular cancer types. RESULTS: We propose a novel Mutated Pathway Visible Neural Network (MPVNN) architecture, designed using prior signaling pathway knowledge and random replacement of known pathway edges using gene mutation data simulating signal flow disruption. As a case study, we use the PI3K-Akt pathway and demonstrate overall improved cancer-specific survival risk prediction of MPVNN over other similar-sized NN and standard survival analysis methods. We show that trained MPVNN architecture interpretation, which points to smaller sets of genes connected by signal flow within the PI3K-Akt pathway that is important in risk prediction for particular cancer types, is reliable. AVAILABILITY AND IMPLEMENTATION: The data and code are available at https://github.com/gourabghoshroy/MPVNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gourab Ghosh Roy, Nicholas Geard, Karin Verspoor, Shan He 0001 |
Bioinform. | 2 |
| 2021 | Managing Trajectories and Interactions During a Pandemic: A Trajectory Similarity-based Approach (Demo Paper)abstractCOVID-19 has brought about substantial social, economic and health related burdens, motivating different control measures from policy makers worldwide. Contact tracing plays a pivotal role in the COVID-19 era. However, contact tracing is by nature entirely retrospective: it can only identify contacts of known or suspected cases. Our proposed system is prospective, aiming to 'create' networks that will ultimately make contact tracing and pandemic management easier. As contact tracing seeks to reconstruct the underlying interaction network, we can improve the process by reducing the complexity of contact network structure; we introduce a method for reducing contact network complexity through strategic scheduling. The method functions through pairwise comparison of individual trajectories in a coordinate space of activities, locations, and time intervals. We demonstrate the method through a simulated scenario where individuals (students) register for activities using a mobile application in a campus. The application then applies our algorithm to provide individuals with schedules that reduce the complexity of the overall network, without compromising individual privacy. Edward Buckland, Egemen Tanin, Nicholas Geard, Cameron Zachreson, Hairuo Xie, Hanan Samet |
SIGSPATIAL/GIS | 3 |
| 2021 | PoLoBag: Polynomial Lasso Bagging for signed gene regulatory network inference from expression dataabstractMOTIVATION: Inferring gene regulatory networks (GRNs) from expression data is a significant systems biology problem. A useful inference algorithm should not only unveil the global structure of the regulatory mechanisms but also the details of regulatory interactions such as edge direction (from regulator to target) and sign (activation/inhibition). Many popular GRN inference algorithms cannot infer edge signs, and those that can infer signed GRNs cannot simultaneously infer edge directions or network cycles. RESULTS: To address these limitations of existing algorithms, we propose Polynomial Lasso Bagging (PoLoBag) for signed GRN inference with both edge directions and network cycles. PoLoBag is an ensemble regression algorithm in a bagging framework where Lasso weights estimated on bootstrap samples are averaged. These bootstrap samples incorporate polynomial features to capture higher-order interactions. Results demonstrate that PoLoBag is consistently more accurate for signed inference than state-of-the-art algorithms on simulated and real-world expression datasets. AVAILABILITY AND IMPLEMENTATION: Algorithm and data are freely available at https://github.com/gourabghoshroy/PoLoBag. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gourab Ghosh Roy, Nicholas Geard, Karin Verspoor, Shan He 0001 |
Bioinform. | 2 |
| 2021 | Automatic consistency assurance for literature-based gene ontology annotationabstractBACKGROUND: Literature-based gene ontology (GO) annotation is a process where expert curators use uniform expressions to describe gene functions reported in research papers, creating computable representations of information about biological systems. Manual assurance of consistency between GO annotations and the associated evidence texts identified by expert curators is reliable but time-consuming, and is infeasible in the context of rapidly growing biological literature. A key challenge is maintaining consistency of existing GO annotations as new studies are published and the GO vocabulary is updated. RESULTS: In this work, we introduce a formalisation of biological database annotation inconsistencies, identifying four distinct types of inconsistency. We propose a novel and efficient method using state-of-the-art text mining models to automatically distinguish between consistent GO annotation and the different types of inconsistent GO annotation. We evaluate this method using a synthetic dataset generated by directed manipulation of instances in an existing corpus, BC4GO. We provide detailed error analysis for demonstrating that the method achieves high precision on more confident predictions. CONCLUSIONS: Two models built using our method for distinct annotation consistency identification tasks achieved high precision and were robust to updates in the GO vocabulary. Our approach demonstrates clear value for human-in-the-loop curation scenarios. Jiyu Chen, Nicholas Geard, Justin Zobel, Karin Verspoor |
BMC Bioinform. | 2 |
| 2020 | Epidemiological consequences of enduring strain-specific immunity requiring repeated episodes of infectionabstractGroup A Streptococcus (GAS) skin infections are caused by a diverse array of strain types and are highly prevalent in disadvantaged populations. The role of strain-specific immunity in preventing GAS infections is poorly understood, representing a critical knowledge gap in vaccine development. A recent GAS murine challenge study showed evidence that sterilising strain-specific and enduring immunity required two skin infections by the same GAS strain within three weeks. This mechanism of developing enduring immunity may be a significant impediment to the accumulation of immunity in populations. We used an agent-based mathematical model of GAS transmission to investigate the epidemiological consequences of enduring strain-specific immunity developing only after two infections with the same strain within a specified interval. Accounting for uncertainty when correlating murine timeframes to humans, we varied this maximum inter-infection interval from 3 to 420 weeks to assess its impact on prevalence and strain diversity, and considered additional scenarios where no maximum inter-infection interval was specified. Model outputs were compared with longitudinal GAS surveillance observations from northern Australia, a region with endemic infection. We also assessed the likely impact of a targeted strain-specific multivalent vaccine in this context. Our model produced patterns of transmission consistent with observations when the maximum inter-infection interval for developing enduring immunity was 19 weeks. Our vaccine analysis suggests that the leading multivalent GAS vaccine may have limited impact on the prevalence of GAS in populations in northern Australia if strain-specific immunity requires repeated episodes of infection. Our results suggest that observed GAS epidemiology from disease endemic settings is consistent with enduring strain-specific immunity being dependent on repeated infections with the same strain, and provide additional motivation for relevant human studies to confirm the human immune response to GAS skin infection. Rebecca H. Chisholm, Nikki Sonenberg, Jake A. Lacey, Malcolm I. McDonald, Manisha Pandey, Mark R. Davies, Steven Y. C. Tong, Jodie McVernon, Nicholas Geard |
PLoS Comput. Biol. | 9 |
| 2016 | Job Insecurity in Academic Research Employment: An Agent-Based Model
Ian B. Wood, Nicholas Geard, Eric Silverman |
ALIFE | 2 |
| 2014 | The Practice of Agent-Based Model VisualizationabstractWe discuss approaches to agent-based model visualization. Agent-based modeling has its own requirements for visualization, some shared with other forms of simulation software, and some unique to this approach. In particular, agent-based models are typified by complexity, dynamism, nonequilibrium and transient behavior, heterogeneity, and a researcher's interest in both individual- and aggregate-level behavior. These are all traits requiring careful consideration in the design, experimentation, and communication of results. In the case of all but final communication for dissemination, researchers may not make their visualizations public. Hence, the knowledge of how to visualize during these earlier stages is unavailable to the research community in a readily accessible form. Here we explore means by which all phases of agent-based modeling can benefit from visualization, and we provide examples from the available literature and online sources to illustrate key stages and techniques. Alan Dorin, Nicholas Geard |
Artif. Life | 2 |
| 2010 | Stability in Flux - Group Dynamics in Adaptive Networks
Nicholas Geard, John Bryden, Sebastian Funk, Vincent A. A. Jansen, Seth Bullock |
ALIFE | 1 |
| 2010 | Adaptive Networks: Theory, Models and Applications. T. Gross and H. Sayama (Eds.). (2009, Springer-Verlag.) GBP108, $159, 332 pages, 162 illustrations, 15 in color (hardcover)
Nicholas Geard |
Artif. Life | 1 |
| 2008 | Group formation and social evolution - a computational model
Nicholas Geard, Seth Bullock |
ALIFE | 1 |
| 2008 | LinMap: Visualizing Complexity Gradients in Evolutionary LandscapesabstractThis article describes an interactive visualization tool, LinMap, for exploring the structure of complexity gradients in evolutionary landscapes. LinMap is a computationally efficient and intuitive tool for visualizing and exploring multidimensional parameter spaces. An artificial cell lineage model is presented that allows complexity to be quantified according to several different developmental and phenotypic metrics. LinMap is applied to the evolutionary landscapes generated by this model to demonstrate that different definitions of complexity produce different gradients across the same landscape; that landscapes are characterized by a phase transition between proliferating and quiescent cell lineages where both complexity and diversity are maximized; and that landscapes defined by adaptive fitness and complexity can display different topographical features. Nicholas Geard, Janet Wiles |
Artif. Life | 1 |
| 2005 | Maximally rugged NK landscapes contain the highest peaksabstractNK models provide a family of tunably rugged fitness landscapes used in a wide range of evolutionary computation studies. It is well known that the average height of local optima regresses to the mean of the landscape with increasing epistasis, k. This fact has been confirmed using both theoretical studies of landscape structure and empirical studies of evolutionary search. We show that the global optimum behaves quite differently: the expected value of the global maximum is highest in the maximally rugged case. Furthermore, we demonstrate that this expected value increases with K, despite the fact that the average fitness of the local optima decreases. That is, the highest peaks are found in the most rugged landscapes, scattered amongst masses of low-lying peaks. We find the asymptotic value of the global optimum as N approaches infinity for both the smooth and maximally rugged cases. In evolutionary search, the optima that are found reflect the local optima that exist in the landscape, the size of these optima -- which corresponds to the size of their basins of attraction, and the effort expended in the search process. Increasing the level of epistasis in an NK landscape stochastically introduces higher peaks, but renders them exponentially more difficult to find. Benjamin Skellett, Benjamin Cairns, Nicholas Geard, Bradley Tonkes, Janet Wiles |
GECCO | 3 |
| 2005 | A Gene Network Model for Developing Cell LineagesabstractBiological development is a remarkably complex process. A single cell, in an appropriate environment, contains sufficient information to generate a variety of differentiated cell types, whose spatial and temporal dynamics interact to form detailed morphological patterns. While several different physical and chemical processes play an important role in the development of an organism, the locus of control is the cell's gene regulatory network. We designed a dynamic recurrent gene network (DRGN) model and evaluated its ability to control the developmental trajectories of cells during embryogenesis. Three tasks were developed to evaluate the model, inspired by cell lineage specification in C. elegans, describing the variation in gene activity required for early cell diversification, combinatorial control of cell lineages, and cell lineage termination. Three corresponding sets of simulations compared performance on the tasks for different gene network sizes, demonstrating the ability of DRGNs to perform the tasks with minimal external input. The model and task definition represent a new means of linking the fundamental properties of genetic networks with the topology of the cell lineages whose development they control. Nicholas Geard, Janet Wiles |
Artif. Life | 1 |
| 2003 | Structure and dynamics of a gene network model incorporating small RNAsabstractAs advances in molecular biology continue to reveal additional layers of complexity in gene regulation, computational models need to incorporate additional features to explore the implications of new theories and hypotheses. It has recently been suggested that eukaryotic organisms owe their phenotypic complexity and diversity to the exploitation of small RNAs as signalling molecules. Previous models of genetic systems are, for several reasons, inadequate to investigate this theory. In this study, we present an artificial genome model of genetic regulatory networks based upon previous work by Torsten Reil, and demonstrate how this model generates networks with biologically plausible structural and dynamic properties. We also extend the model to explore the implications of incorporating regulation by small RNA molecules in a gene network. We demonstrate how, using these signals, highly connected networks can display dynamics that are more stable than expected given their level of connectivity. Nicholas Geard, Janet Wiles |
IEEE Congress on Evolutionary Computation | 1 |
| 2002 | Diversity maintenance on neutral landscapes: an argument for recombinationabstractIt has been demonstrated that several standard evolutionary computation test problems can be solved by a simple hill climbing search algorithm-often more efficiently than by a population based evolutionary algorithm. There remain some classes of problems, however, for which maintaining a genetically diverse population is essential in order to discover the optimal solution. In biological populations, diversity maintenance is important to enable populations to adapt to rapidly changing environments and to exploit environmental niches. We demonstrate that on a neutral landscape recombination allows a population to maintain a significantly greater level of genetic diversity through the transition between two fitness layers. Recombination may therefore have a role to play in maintaining population diversity across fitness transitions. Nicholas Geard, Janet Wiles |
IEEE Congress on Evolutionary Computation | 1 |
| 2002 | A comparison of neutral landscapes - NK, NKp and NKqabstractRecent research in molecular evolution has raised awareness of the importance of selective neutrality. Several different models of neutrality have been proposed based on Kauffman's well-known NK landscape model. Two of these models, NKp and NKq, are investigated and found to display significantly different structural properties. The fitness distributions of these neutral landscapes reveal that their levels of correlation with non-neutral landscapes are significantly different, as are the distributions of neutral mutations. In this paper we describe a series of simulations of a hill climbing search algorithm on NK, NKp and NKq landscapes with varying levels of epistatic interaction. These simulations demonstrate differences in the way that epistatic interaction affects the 'searchability' of neutral landscapes. We conclude that the method used to implement neutrality has an impact on both the structure of the resulting landscapes and on the performance of evolutionary search algorithms on these landscapes. These model-dependent effects must be taken into consideration when modelling biological phenomena. Nicholas Geard, Janet Wiles, Jennifer Hallinan, Bradley Tonkes, Benjamin Skellett |
IEEE Congress on Evolutionary Computation | 1 |