David R. Gilbert

dblp:g/DavidRGilbert · also David Roger Gilbert · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorTheory of computation · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorComputer networks · 2 · 2 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Insensitive Games: Game Semantics for Modal Insensitivity
Can Baskent, David R. Gilbert, Giorgio Venturi
WoLLIC2
2025 Logics of spatial isolation
abstract
Abstract In the vein of recent work that provides non-normal modal interpretations of various topological operators, this paper proposes a modal logic for a spatial isolation operator. Focusing initially on neighbourhood systems, we prove several characterization results, demonstrating the adequacy of the interpretation and highlighting certain semantic insensitivities that result from the relative expressive weakness of the isolation operator. We then transition to the topological setting, proving a topological characterization result.
Can Baskent, David R. Gilbert, Giorgio Venturi
J. Log. Comput.2
2024 A Logic of Isolation
Can Baskent, David R. Gilbert, Giorgio Venturi
WoLLIC2
2022 Hybrid modelling of biological systems: current progress and future prospects
abstract
Integrated modelling of biological systems is becoming a necessity for constructing models containing the major biochemical processes of such systems in order to obtain a holistic understanding of their dynamics and to elucidate emergent behaviours. Hybrid modelling methods are crucial to achieve integrated modelling of biological systems. This paper reviews currently popular hybrid modelling methods, developed for systems biology, mainly revealing why they are proposed, how they are formed from single modelling formalisms and how to simulate them. By doing this, we identify future research requirements regarding hybrid approaches for further promoting integrated modelling of biological systems.
Fei Liu 0006, Monika Heiner, David R. Gilbert
Briefings Bioinform.3
2021 Hybrid modelling of biological systems using fuzzy continuous Petri nets
abstract
Integrated modelling of biological systems is challenged by composing components with sufficient kinetic data and components with insufficient kinetic data or components built only using experts' experience and knowledge. Fuzzy continuous Petri nets (FCPNs) combine continuous Petri nets with fuzzy inference systems, and thus offer an hybrid uncertain/certain approach to integrated modelling of such biological systems with uncertainties. In this paper, we give a formal definition and a corresponding simulation algorithm of FCPNs, and briefly introduce the FCPN tool that we have developed for implementing FCPNs. We then present a methodology and workflow utilizing FCPNs to achieve hybrid (uncertain/certain) modelling of biological systems illustrated with a case study of the Mercaptopurine metabolic pathway. We hope this research will promote the wider application of FCPNs and address the uncertain/certain integrated modelling challenge in the systems biology area.
Fei Liu 0006, Wujie Sun, Monika Heiner, David R. Gilbert
Briefings Bioinform.4
2020 Fuzzy Petri nets for modelling of uncertain biological systems
abstract
The modelling of biological systems is accompanied with epistemic uncertainties that range from structural uncertainty to parametric uncertainty due to such limitations as insufficient understanding of the underlying mechanism and incomplete measurement data of a system. Fuzzy logic approaches such as fuzzy Petri nets (FPNs) are effective in addressing these issues. In this paper, we review FPNs that have been used for modelling uncertain biological systems, which we classify in three categories: basic fuzzy Petri nets, fuzzy quantitative Petri nets and Petri nets with fuzzy kinetic parameters. For each category of these FPNs, we summarize its modelling capabilities and current applications, discuss its merits and drawbacks and give suggestions for further research. This understanding on how to use FPNs for modelling uncertain biological systems will assist readers in selecting appropriate FPN classes for specific modelling circumstances. This review may also promote the extensive research and application of FPNs in the systems biology area.
Fei Liu 0006, Monika Heiner, David R. Gilbert
Briefings Bioinform.3
2019 Towards dynamic genome-scale models
abstract
The analysis of the dynamic behaviour of genome-scale models of metabolism (GEMs) currently presents considerable challenges because of the difficulties of simulating such large and complex networks. Bacterial GEMs can comprise about 5000 reactions and metabolites, and encode a huge variety of growth conditions; such models cannot be used without sophisticated tool support. This article is intended to aid modellers, both specialist and non-specialist in computerized methods, to identify and apply a suitable combination of tools for the dynamic behaviour analysis of large-scale metabolic designs. We describe a methodology and related workflow based on publicly available tools to profile and analyse whole-genome-scale biochemical models. We use an efficient approximative stochastic simulation method to overcome problems associated with the dynamic simulation of GEMs. In addition, we apply simulative model checking using temporal logic property libraries, clustering and data analysis, over time series of reaction rates and metabolite concentrations. We extend this to consider the evolution of reaction-oriented properties of subnets over time, including dead subnets and functional subsystems. This enables the generation of abstract views of the behaviour of these models, which can be large-up to whole genome in size-and therefore impractical to analyse informally by eye. We demonstrate our methodology by applying it to a reduced model of the whole-genome metabolism of Escherichia coli K-12 under different growth conditions. The overall context of our work is in the area of model-based design methods for metabolic engineering and synthetic biology.
David R. Gilbert, Monika Heiner, Yasoda Jayaweera, Christian Rohr
Briefings Bioinform.1
2019 Coloured Petri nets for multilevel, multiscale and multidimensional modelling of biological systems
abstract
Owing to the availability of data of one biological phenomenon at different levels/scales, modelling of biological systems is moving from single level/scale to multiple levels/scales, which introduces a number of challenges. Coloured Petri nets (ColPNs) have been successfully applied to multilevel, multiscale and multidimensional modelling of some biological systems, addressing many of these challenges. In this article, we first review the basics of ColPNs and some popular extensions, and then their applications for multilevel, multiscale and multidimensional modelling of biological systems. This understanding of how to use ColPNs for modelling biological systems will assist readers in selecting appropriate ColPN classes for specific modelling circumstances.
Fei Liu 0006, Monika Heiner, David R. Gilbert
Briefings Bioinform.3
2019 Spatial quorum sensing modelling using coloured hybrid Petri nets and simulative model checking
abstract
BACKGROUND: Quorum sensing drives biofilm formation in bacteria in order to ensure that biofilm formation only occurs when colonies are of a sufficient size and density. This spatial behaviour is achieved by the broadcast communication of an autoinducer in a diffusion scenario. This is of interest, for example, when considering the role of gut microbiota in gut health. This behaviour occurs within the context of the four phases of bacterial growth, specifically in the exponential stage (phase 2) for autoinducer production and the stationary stage (phase 3) for biofilm formation. RESULTS: We have used coloured hybrid Petri nets to step-wise develop a flexible computational model for E.coli biofilm formation driven by Autoinducer 2 (AI-2) which is easy to configure for different notions of space. The model describes the essential components of gene transcription, signal transduction, extra and intra cellular transport, as well as the two-phase nature of the system. We build on a previously published non-spatial stochastic Petri net model of AI-2 production, keeping the assumptions of a limited nutritional environment, and our spatial hybrid Petri net model of biofilm formation, first presented at the NETTAB 2017 workshop. First we consider the two models separately without space, and then combined, and finally we add space. We describe in detail our step-wise model development and validation. Our simulation results support the expected behaviour that biofilm formation is increased in areas of higher bacterial colony size and density. Our analysis techniques include behaviour checking based on linear time temporal logic. CONCLUSIONS: The advantages of our modelling and analysis approach are the description of quorum sensing and associated biofilm formation over two phases of bacterial growth, taking into account bacterial spatial distribution using a flexible and easy to maintain computational model. All computational results are reproducible.
David R. Gilbert, Monika Heiner, Leila Ghanbar, Jacek Chodak
BMC Bioinform.1
2018 Emerging ensembles of kinetic parameters to characterize observed metabolic phenotypes
abstract
BACKGROUND: Determining the value of kinetic constants for a metabolic system in the exact physiological conditions is an extremely hard task. However, this kind of information is of pivotal relevance to effectively simulate a biological phenomenon as complex as metabolism. RESULTS: To overcome this issue, we propose to investigate emerging properties of ensembles of sets of kinetic constants leading to the biological readout observed in different experimental conditions. To this aim, we exploit information retrievable from constraint-based analyses (i.e. metabolic flux distributions at steady state) with the goal to generate feasible values for kinetic constants exploiting the mass action law. The sets retrieved from the previous step will be used to parametrize a mechanistic model whose simulation will be performed to reconstruct the dynamics of the system (until reaching the metabolic steady state) for each experimental condition. Every parametrization that is in accordance with the expected metabolic phenotype is collected in an ensemble whose features are analyzed to determine the emergence of properties of a phenotype. In this work we apply the proposed approach to identify ensembles of kinetic parameters for five metabolic phenotypes of E. Coli, by analyzing five different experimental conditions associated with the ECC2comp model recently published by Hädicke and collaborators. CONCLUSIONS: Our results suggest that the parameter values of just few reactions are responsible for the emergence of a metabolic phenotype. Notably, in contrast with constraint-based approaches such as Flux Balance Analysis, the methodology used in this paper does not require to assume that metabolism is optimizing towards a specific goal.
Riccardo Colombo, Chiara Damiani, David R. Gilbert, Monika Heiner, Giancarlo Mauri, Dario Pescini
BMC Bioinform.3
2018 Petri-net-based 2D design of DNA walker circuits
abstract
We consider localised DNA computation, where a DNA strand walks along a binary decision graph to compute a binary function. One of the challenges for the design of reliable walker circuits consists in leakage transitions, which occur when a walker jumps into another branch of the decision graph. We automatically identify leakage transitions, which allows for a detailed qualitative and quantitative assessment of circuit designs, design comparison, and design optimisation. The ability to identify leakage transitions is an important step in the process of optimising DNA circuit layouts where the aim is to minimise the computational error inherent in a circuit while minimising the area of the circuit. Our 2D modelling approach of DNA walker circuits relies on coloured stochastic Petri nets which enable functionality, topology and dimensionality all to be integrated in one two-dimensional model. Our modelling and analysis approach can be easily extended to 3-dimensional walker systems.
David R. Gilbert, Monika Heiner, Christian Rohr
Nat. Comput.1
2015 Computational models for inferring biochemical networks
Silvia Rausanu, Crina Grosan, Zujian Wu, Ovidiu Parvu, Ramona Stoica, David R. Gilbert
Neural Comput. Appl.6
2015 Advances in Computational Methods in Systems Biology
David R. Gilbert, Monika Heiner
Theor. Comput. Sci.1
2014 Speeding up systems biology simulations of biochemical pathways using condor
abstract
SUMMARY Systems biology is a scientific field that uses computational modelling to study biological and biochemical systems. The simulation and analysis of models of these systems typically explore behaviour over a wide range of parameter values; as such, they are usually characterised by the need for nontrivial amounts of computing power. Grid computing provides access to such computational resources. In previous research, we created the grid‐enabledbiochemical networks simulation environmentto attempt to speed up system biology simulations over a grid (the UK National Grid Service and ScotGrid). Following on from this work, we have created thesimulation modelling of the epidermal growth factor receptor microtubule‐associated proteinkinase pathway utility, a standalone simulation tool dedicated to the modelling and analysis of theepidermal growth factor receptor microtubule‐associated protein kinase pathway. This builds on experiences frombiochemical networks simulation environmentby decoupling the simulation modelling elements from the Grid middleware. This new utility enables us to interface with different grid technologies. This paper therefore describes the new SIMAP utility and an empirical investigation of its performance when deployed over a desktop grid based on the high throughput computing middlewareCondor. We present our results based on a case study with a model of themammalian ErbB signalling pathway, a pathway strongly linked to cancer. Copyright © 2013 John Wiley & Sons, Ltd.
Simon J. E. Taylor, Navonil Mustafee, Qian Gao 0001, David R. Gilbert
Concurr. Comput. Pract. Exp.6
2013 Colouring Space - A Coloured Framework for Spatial Modelling in Systems Biology
David R. Gilbert, Monika Heiner, Fei Liu 0006, Nigel J. Saunders
Petri Nets1
2013 Evolving biochemical systems
abstract
The interaction of biological compounds in cells has been enforced to a proper understanding by the numerous bioinformatics projects which contributed with a vast amount of biological information. The construction of biochemical systems (systems of chemical reactions) which include both topology and kinetic rates of the chemical reactions is an NP-hard problem. In this paper we propose a hybrid architecture which combines genetic programming and simulated annealing in order to generate and optimize both the topology (the network) and the reaction rates of a biochemical systems. Simulations and analysis of two real models show promising results for the proposed method.
Silvia Rausanu, Crina Grosan, Zujian Wu, Ovidiu Parvu, David R. Gilbert
IEEE Congress on Evolutionary Computation5
2013 Handling Uncertainty in Dynamic Models: The Pentose Phosphate Pathway in Trypanosoma brucei
abstract
Dynamic models of metabolism can be useful in identifying potential drug targets, especially in unicellular organisms. A model of glycolysis in the causative agent of human African trypanosomiasis, Trypanosoma brucei, has already shown the utility of this approach. Here we add the pentose phosphate pathway (PPP) of T. brucei to the glycolytic model. The PPP is localized to both the cytosol and the glycosome and adding it to the glycolytic model without further adjustments leads to a draining of the essential bound-phosphate moiety within the glycosome. This phosphate "leak" must be resolved for the model to be a reasonable representation of parasite physiology. Two main types of theoretical solution to the problem could be identified: (i) including additional enzymatic reactions in the glycosome, or (ii) adding a mechanism to transfer bound phosphates between cytosol and glycosome. One example of the first type of solution would be the presence of a glycosomal ribokinase to regenerate ATP from ribose 5-phosphate and ADP. Experimental characterization of ribokinase in T. brucei showed that very low enzyme levels are sufficient for parasite survival, indicating that other mechanisms are required in controlling the phosphate leak. Examples of the second type would involve the presence of an ATP:ADP exchanger or recently described permeability pores in the glycosomal membrane, although the current absence of identified genes encoding such molecules impedes experimental testing by genetic manipulation. Confronted with this uncertainty, we present a modeling strategy that identifies robust predictions in the context of incomplete system characterization. We illustrate this strategy by exploring the mechanism underlying the essential function of one of the PPP enzymes, and validate it by confirming the model predictions experimentally.
Eduard J. Kerkhoven, Fiona Achcar, Vincent P. Alibu, Richard J. Burchmore, Ian H. Gilbert, Maciej Trybilo, Nicole N. Driessen, David R. Gilbert, Rainer Breitling, Barbara M. Bakker, Michael P. Barrett
PLoS Comput. Biol.8
2013 Multiscale Modeling and Analysis of Planar Cell Polarity in the Drosophila Wing
abstract
Modeling across multiple scales is a current challenge in Systems Biology, especially when applied to multicellular organisms. In this paper, we present an approach to model at different spatial scales, using the new concept of Hierarchically Colored Petri Nets (HCPN). We apply HCPN to model a tissue comprising multiple cells hexagonally packed in a honeycomb formation in order to describe the phenomenon of Planar Cell Polarity (PCP) signaling in Drosophila wing. We have constructed a family of related models, permitting different hypotheses to be explored regarding the mechanisms underlying PCP. In addition our models include the effect of well-studied genetic mutations. We have applied a set of analytical techniques including clustering and model checking over time series of primary and secondary data. Our models support the interpretation of biological observations reported in the literature.
Qian Gao 0001, David R. Gilbert, Monika Heiner, Fei Liu 0006, Daniele Maccagnola, David Tree
IEEE ACM Trans. Comput. Biol. Bioinform.2
2012 A Hybrid Approach to Piecewise Modelling of Biochemical Systems
Zujian Wu, Shengxiang Yang, David R. Gilbert
PPSN (1)3
2011 How Might Petri Nets Enhance Your Systems Biology Toolkit
Monika Heiner, David R. Gilbert
Petri Nets2
2011 Bootstrapping Parameter Estimation in Dynamic Systems
Huma Lodhi, David R. Gilbert
Discovery Science2
2010 An optimized TOPS+ comparison method for enhanced TOPS models
abstract
BACKGROUND: Although methods based on highly abstract descriptions of protein structures, such as VAST and TOPS, can perform very fast protein structure comparison, the results can lack a high degree of biological significance. Previously we have discussed the basic mechanisms of our novel method for structure comparison based on our TOPS+ model (Topological descriptions of Protein Structures Enhanced with Ligand Information). In this paper we show how these results can be significantly improved using parameter optimization, and we call the resulting optimised TOPS+ method as advanced TOPS+ comparison method i.e. advTOPS+. RESULTS: We have developed a TOPS+ string model as an improvement to the TOPS 123 graph model by considering loops as secondary structure elements (SSEs) in addition to helices and strands, representing ligands as first class objects, and describing interactions between SSEs, and SSEs and ligands, by incoming and outgoing arcs, annotating SSEs with the interaction direction and type. Benchmarking results of an all-against-all pairwise comparison using a large dataset of 2,620 non-redundant structures from the PDB40 dataset 4 demonstrate the biological significance, in terms of SCOP classification at the superfamily level, of our TOPS+ comparison method. CONCLUSIONS: Our advanced TOPS+ comparison shows better performance on the PDB40 dataset 4 compared to our basic TOPS+ method, giving 90% accuracy for SCOP alpha+beta; a 6% increase in accuracy compared to the TOPS and basic TOPS+ methods. It also outperforms the TOPS, basic TOPS+ and SSAP comparison methods on the Chew-Kedem dataset 5, achieving 98% accuracy. SOFTWARE AVAILABILITY: The TOPS+ comparison server is available at http://balabio.dcs.gla.ac.uk/mallika/WebTOPS/.
Mallika Veeramalai, David R. Gilbert, Gabriel Valiente
BMC Bioinform.2
2009 Prediction of protein-protein interaction types using association rule based classification
abstract
BACKGROUND: Protein-protein interactions (PPI) can be classified according to their characteristics into, for example obligate or transient interactions. The identification and characterization of these PPI types may help in the functional annotation of new protein complexes and in the prediction of protein interaction partners by knowledge driven approaches. RESULTS: This work addresses pattern discovery of the interaction sites for four different interaction types to characterize and uses them for the prediction of PPI types employing Association Rule Based Classification (ARBC) which includes association rule generation and posterior classification. We incorporated domain information from protein complexes in SCOP proteins and identified 354 domain-interaction sites. 14 interface properties were calculated from amino acid and secondary structure composition and then used to generate a set of association rules characterizing these domain-interaction sites employing the APRIORI algorithm. Our results regarding the classification of PPI types based on a set of discovered association rules shows that the discriminative ability of association rules can significantly impact on the prediction power of classification models. We also showed that the accuracy of the classification can be improved through the use of structural domain information and also the use of secondary structure content. CONCLUSION: The advantage of our approach is that we can extract biologically significant information from the interpretation of the discovered association rules in terms of understandability and interpretability of rules. A web application based on our method can be found at http://bioinfo.ssu.ac.kr/~shpark/picasso/
Sung-Hee Park, José A. Reyes, David R. Gilbert, Ji Woong Kim, Sangsoo Kim
BMC Bioinform.3
2008 A structured approach for the engineering of biochemical network models, illustrated for signalling pathways
abstract
Quantitative models of biochemical networks (signal transduction cascades, metabolic pathways, gene regulatory circuits) are a central component of modern systems biology. Building and managing these complex models is a major challenge that can benefit from the application of formal methods adopted from theoretical computing science. Here we provide a general introduction to the field of formal modelling, which emphasizes the intuitive biochemical basis of the modelling process, but is also accessible for an audience with a background in computing science and/or model engineering. We show how signal transduction cascades can be modelled in a modular fashion, using both a qualitative approach--qualitative Petri nets, and quantitative approaches--continuous Petri nets and ordinary differential equations (ODEs). We review the major elementary building blocks of a cellular signalling model, discuss which critical design decisions have to be made during model building, and present a number of novel computational tools that can help to explore alternative modular models in an easy and intuitive manner. These tools, which are based on Petri net theory, offer convenient ways of composing hierarchical ODE models, and permit a qualitative analysis of their behaviour. We illustrate the central concepts using signal transduction as our main example. The ultimate aim is to introduce a general approach that provides the foundations for a structured formal engineering of large-scale models of biochemical networks.
Rainer Breitling, David R. Gilbert, Monika Heiner, Richard J. Orton
Briefings Bioinform.2
2008 MetaNetter: inference and visualization of high-resolution metabolomic networks
abstract
Abstract Summary: We present a Cytoscape plugin for the inference and visualization of networks from high-resolution mass spectrometry metabolomic data. The software also provides access to basic topological analysis. This open source, multi-platform software has been successfully used to interpret metabolomic experiments and will enable others using filtered, high mass accuracy mass spectrometric data sets to build and analyse networks. Availability: http://compbio.dcs.gla.ac.uk/fabien/abinitio/abinitio.html Contact: [email protected] Supplementary information: http://compbio.dcs.gla.ac.uk/fabien/abinitio/doc/Supplementary.pdf
Fabien Jourdan, Rainer Breitling, Michael P. Barrett, David R. Gilbert
Bioinform.4
2008 A novel method for comparing topological models of protein structures enhanced with ligand information
abstract
UNLABELLED: We introduce TOPS+ strings, a highly abstract string-based model of protein topology that permits efficient computation of structure comparison, and can optionally represent ligand information. In this model, we consider loops as secondary structure elements (SSEs) as well as helices and strands; in addition we represent ligands as first class objects. Interactions between SSEs and between SSEs and ligands are described by incoming/outgoing arcs and ligand arcs, respectively; and SSEs are annotated with arc interaction direction and type. We are able to abstract away from the ligands themselves, to give a model characterized by a regular grammar rather than the context sensitive grammar of the original TOPS model. Our TOPS+ strings model is sufficiently descriptive to obtain biologically meaningful results and has the advantage of permitting fast string-based structure matching and comparison as well as avoiding issues of Non-deterministic Polynomial time (NP)-completeness associated with graph problems. Our structure comparison method is computationally more efficient in identifying distantly related proteins than BLAST, CLUSTALW, SSAP and TOPS because of the compact and abstract string-based representation of protein structure which records both topological and biochemical information including the functionally important loop regions of the protein structures. The accuracy of our comparison method is comparable with that of TOPS. Also, we have demonstrated that our TOPS+ strings method out-performs the TOPS method for the ligand-dependent protein structures and provides biologically meaningful results. AVAILABILITY: The TOPS+ strings comparison server is available from http://balabio.dcs.gla.ac.uk/mallika/WebTOPS/topsplus.html.
Mallika Veeramalai, David R. Gilbert
Bioinform.2
2007 Fast Structural Similarity Search Based on Topology String Matching
Sung-Hee Park, David R. Gilbert, Keun Ho Ryu
APBC2
2007 Discriminating Microbial Species Using Protein Sequence Properties and Machine Learning
Ali Al-Shahib, David R. Gilbert, Rainer Breitling
IDEAL2
2007 Assessment of the probabilities for evolutionary structural changes in protein folds
abstract
MOTIVATION: The evolution of protein sequences can be described by a stepwise process, where each step involves changes of a few amino acids. In a similar manner, the evolution of protein folds can be at least partially described by an analogous process, where each step involves comparatively simple changes affecting few secondary structure elements. A number of such evolution steps, justified by biologically confirmed examples, have previously been proposed by other researchers. However, unlike the situation with sequences, as far as we know there have been no attempts to estimate the comparative probabilities for different kinds of such structural changes. RESULTS: We have tried to assess the comparative probabilities for a number of known structural changes, and to relate the probabilities of such changes with the distance between protein sequences. We have formalized these structural changes using a topological representation of structures (TOPS), and have developed an algorithm for measuring structural distances that involve few evolutionary steps. The probabilities of structural changes then were estimated on the basis of all-against-all comparisons of the sequence and structure of protein domains from the CATH-95 representative set. The results obtained are reasonably consistent for a number of different data subsets and permit the identification of several 'most popular' types of evolutionary changes in protein structure. The results also suggest that alterations in protein structure are more likely to occur when the sequence similarity is >10% (the average similarity being approximately 6% for the data sets employed in this study), and that the distribution of probabilities of structural changes is fairly uniform within the interval of 15-50% sequence similarity. AVAILABILITY: The algorithms have been implemented on the Windows operating system in C++ and using the Borland Visual Component Library. The source code is available on request from the first author. The data sets used for this study (representative sets of protein domains, matrices of sequence similarities and structural distances) are available on http://bioinf.mii.lu.lv/epsrc_project/struct_ev.html.
Juris Viksna, David R. Gilbert
Bioinform.2
2007 Automatic generation of 3D motifs for classification of protein binding sites
abstract
BACKGROUND: Since many of the new protein structures delivered by high-throughput processes do not have any known function, there is a need for structure-based prediction of protein function. Protein 3D structures can be clustered according to their fold or secondary structures to produce classes of some functional significance. A recent alternative has been to detect specific 3D motifs which are often associated to active sites. Unfortunately, there are very few known 3D motifs, which are usually the result of a manual process, compared to the number of sequential motifs already known. In this paper, we report a method to automatically generate 3D motifs of protein structure binding sites based on consensus atom positions and evaluate it on a set of adenine based ligands. RESULTS: Our new approach was validated by generating automatically 3D patterns for the main adenine based ligands, i.e. AMP, ADP and ATP. Out of the 18 detected patterns, only one, the ADP4 pattern, is not associated with well defined structural patterns. Moreover, most of the patterns could be classified as binding site 3D motifs. Literature research revealed that the ADP4 pattern actually corresponds to structural features which show complex evolutionary links between ligases and transferases. Therefore, all of the generated patterns prove to be meaningful. Each pattern was used to query all PDB proteins which bind either purine based or guanine based ligands, in order to evaluate the classification and annotation properties of the pattern. Overall, our 3D patterns matched 31% of proteins with adenine based ligands and 95.5% of them were classified correctly. CONCLUSION: A new metric has been introduced allowing the classification of proteins according to the similarity of atomic environment of binding sites, and a methodology has been developed to automatically produce 3D patterns from that classification. A study of proteins binding adenine based ligands showed that these 3D patterns are not only biochemically meaningful, but can be used for protein classification and annotation.
Jean-Christophe Nebel, Pawel Herzyk, David R. Gilbert
BMC Bioinform.3
2006 Computational methodologies for modelling, analysis and simulation of signalling networks
abstract
This article is a critical review of computational techniques used to model, analyse and simulate signalling networks. We propose a conceptual framework, and discuss the role of signalling networks in three major areas: signal transduction, cellular rhythms and cell-to-cell communication. In order to avoid an overly abstract and general discussion, we focus on three case studies in the areas of receptor signalling and kinase cascades, cell-cycle regulation and wound healing. We report on a variety of modelling techniques and associated tools, in addition to the traditional approach based on ordinary differential equations (ODEs), which provide a range of descriptive and analytical powers. As the field matures, we expect a wider uptake of these alternative approaches for several reasons, including the need to take into account low protein copy numbers and noise and the great complexity of cellular organisation. An advantage offered by many of these alternative techniques, which have their origins in computing science, is the ability to perform sophisticated model analysis which can better relate predicted behaviour and observations.
David R. Gilbert, Hendrik Fuß, Richard J. Orton, Steve Robinson, Vladislav Vyshemirsky, Mary Jo Kurth, C. Stephen Downes, Werner Dubitzky
Briefings Bioinform.1
2006 A lock-and-key model for protein-protein interactions
abstract
MOTIVATION: Protein-protein interaction networks are one of the major post-genomic data sources available to molecular biologists. They provide a comprehensive view of the global interaction structure of an organism's proteome, as well as detailed information on specific interactions. Here we suggest a physical model of protein interactions that can be used to extract additional information at an intermediate level: It enables us to identify proteins which share biological interaction motifs, and also to identify potentially missing or spurious interactions. RESULTS: Our new graph model explains observed interactions between proteins by an underlying interaction of complementary binding domains (lock-and-key model). This leads to a novel graph-theoretical algorithm to identify bipartite subgraphs within protein-protein interaction networks where the underlying data are taken from yeast two-hybrid experimental results. By testing on synthetic data, we demonstrate that under certain modelling assumptions, the algorithm will return correct domain information about each protein in the network. Tests on data from various model organisms show that the local and global patterns predicted by the model are indeed found in experimental data. Using functional and protein structure annotations, we show that bipartite subnetworks can be identified that correspond to biologically relevant interaction motifs. Some of these are novel and we discuss an example involving SH3 domains from the Saccharomyces cerevisiae interactome. AVAILABILITY: The algorithm (in Matlab format) is available (see http://www.maths.strath.ac.uk/~aas96106/lock_key.html).
Julie L. Morrison, Rainer Breitling, Desmond J. Higham, David R. Gilbert
Bioinform.4
2005 Protein structure topological comparison, discovery and matching service
abstract
UNLABELLED: We describe a fold level fast protein comparison and motif matching facility based on the TOPS representation of structure. This provides an update to a previous service at the EBI, with a better graph matching with faster results and visualization of both the structures being compared against and the common pattern of each with the target domain. AVAILABILITY: Web service at http://balabio.dcs.gla.ac.uk/tops or via the main TOPS site at http://www.tops.leeds.ac.uk. Software is also available for download from these sites.
Gilleain M. Torrance, David R. Gilbert, Ioannis Michalopoulos, David R. Westhead
Bioinform.2
2005 GeneRank: Using search engine technology for the analysis of microarray experiments
abstract
BACKGROUND: Interpretation of simple microarray experiments is usually based on the fold-change of gene expression between a reference and a "treated" sample where the treatment can be of many types from drug exposure to genetic variation. Interpretation of the results usually combines lists of differentially expressed genes with previous knowledge about their biological function. Here we evaluate a method--based on the PageRank algorithm employed by the popular search engine Google--that tries to automate some of this procedure to generate prioritized gene lists by exploiting biological background information. RESULTS: GeneRank is an intuitive modification of PageRank that maintains many of its mathematical properties. It combines gene expression information with a network structure derived from gene annotations (gene ontologies) or expression profile correlations. Using both simulated and real data we find that the algorithm offers an improved ranking of genes compared to pure expression change rankings. CONCLUSION: Our modification of the PageRank algorithm provides an alternative method of evaluating microarray experimental results which combines prior knowledge about the underlying network. GeneRank offers an improvement compared to assessing the importance of a gene based on its experimentally observed fold-change alone and may be used as a basis for further analytical developments.
Julie L. Morrison, Rainer Breitling, Desmond J. Higham, David R. Gilbert
BMC Bioinform.4
2005 Franksum: new feature selection method for protein function prediction
abstract
In the study of in silico functional genomics, improving the performance of protein function prediction is the ultimate goal for identifying proteins associated with defined cellular functions. The classical prediction approach is to employ pairwise sequence alignments. However this method often faces difficulties when no statistically significant homologous sequences are identified. An alternative way is to predict protein function from sequence-derived features using machine learning. In this case the choice of possible features which can be derived from the sequence is of vital importance to ensure adequate discrimination to predict function. In this paper we have successfully selected biologically significant features for protein function prediction. This was performed using a new feature selection method (FrankSum) that avoids data distribution assumptions, uses a data independent measurement (p-value) within the feature, identifies redundancy between features and uses an appropriate ranking criterion for feature selection. We have shown that classifiers generated from features selected by FrankSum outperforms classifiers generated from full feature sets, randomly selected features and features selected from the Wrapper method. We have also shown the features are concordant across all species and top ranking features are biologically informative. We conclude that feature selection is vital for successful protein function prediction and FrankSum is one of the feature selection methods that can be applied successfully to such a domain.
Ali Al-Shahib, Rainer Breitling, David R. Gilbert
Int. J. Neural Syst.3
2005 Fast similarity search for protein 3d structures using topological pattern matching based on spatial relations
abstract
Similarity search for protein 3D structures become complex and computationally expensive due to the fact that the size of protein structure databases continues to grow tremendously. Recently, fast structural similarity search systems have been required to put them into practical use in protein structure classification whilst existing comparison systems do not provide comparison results on time. Our approach uses multi-step processing that composes of a preprocessing step to represent geometry of protein structures with spatial objects, a filter step to generate a small candidate set using approximate topological string matching, and a refinement step to compute a structural alignment. This paper describes the preprocessing and filtering for fast similarity search using the discovery of topological patterns of secondary structure elements based on spatial relations. Our system is fully implemented by using Oracle 8i spatial. We have previously shown that our approach has the advantage of speed of performance compared with other approach such as DALI. This work shows that the discovery of topological relations of secondary structure elements in protein structures by using spatial relations of spatial databases is practical for fast structural similarity search for proteins.
Sung-Hee Park, Keun Ho Ryu, David R. Gilbert
Int. J. Neural Syst.3
2004 An Assessment of Feature Relevance in Predicting Protein Function from Sequence
Ali Al-Shahib, Aik Choon Tan, Mark A. Girolami, David R. Gilbert
IDEAL5
2003 An Empirical Comparison of Supervised Machine Learning Techniques in Bioinformatics
Aik Choon Tan, David R. Gilbert
APBC2
2003 An overview of data models for the analysis of biochemical pathways
abstract
Biochemical pathways such as metabolic, regulatory or signal transduction pathways can be viewed as interconnected processes forming an intricate network of functional and physical interactions between molecular species in the cell. The amount of information available on such pathways for different organisms is increasing very rapidly. This is offering the possibility of performing various analyses on the structure of the full network of pathways for one organism as well as across different organisms, and has therefore generated interest in developing databases for storing and managing this information. Analysing these networks remains far from straightforward owing to the nature of the databases, which are often heterogeneous, incomplete or inconsistent. Pathway analysis is hence a challenging problem in systems biology and in bioinformatics. Various forms of data models have been devised for the analysis of biochemical pathways. This paper presents an overview of the types of models used for this purpose, concentrating on those concerned with the structural aspects of biochemical networks. In particular, the different types of data models found in the literature are classified using a unified framework. In addition, how these models have been used in the analysis of biochemical networks is described. This enables us to underline the strengths and weaknesses of the different approaches, as well as to highlight relevant future research directions.
Yves Deville, David R. Gilbert, Jacques van Helden, Shoshana J. Wodak
Briefings Bioinform.2
2001 Multi-agent Systems as Concurrent Constraint Processes
Lubos Brim, David R. Gilbert, Jean-Marie Jacquet, Mojmír Kretínský
SOFSEM2
2001 Pattern Matching and Pattern Discovery Algorithms for Protein Topologies
Juris Viksna, David R. Gilbert
WABI2
2001 Approaches to visualisation in bioinformatics: from dendrograms to Space Explorer
Michael Schroeder 0001, David R. Gilbert, Jacques van Helden, Penny Noy
Inf. Sci.2
2000 FURY: Fuzzy Unification and Resolution Based on Edit Distance
abstract
The authors present a theoretically founded framework for fuzzy unification and resolution based on edit distance over trees. Their framework extends classical unification and resolution conservatively. They prove important properties of the framework and develop the FURY system, which implements the framework efficiently using dynamic programming. The authors evaluate the framework and system on a large problem in the bioinformatics domain, that of detecting typographical errors in an enzyme name database.
David R. Gilbert, Michael Schroeder 0001
BIBE1
1999 Motif-based searching in TOPS protein topology databases
abstract
MOTIVATION: TOPS cartoons are a schematic ion of protein three-dimensional structures in two dimensions, and are used for understanding and manual comparison of protein folds. Recently, an algorithm that produces the cartoons automatically from protein structures has been devised and cartoons have been generated to represent all the structures in the structural databank. There is now a need to be able to define target topological patterns and to search the database for matching domains. RESULTS: We have devised a formal language for describing TOPS diagrams and patterns, and have designed an efficient algorithm to match a pattern to a set of diagrams. A pattern-matching system has been implemented, and tested on a database derived from all the current entries in the Protein Data Bank (15,000 domains). Users can search on patterns selected from a library of motifs or, alternatively, they can define their own search patterns. AVAILABILITY: The system is accessible over the Web at http://tops.ebi.ac.uk/tops
David R. Gilbert, David R. Westhead, Nozomi Nagano, Janet M. Thornton
Bioinform.1
1996 Transformations Between HCLP and PCSP
Michael Jampel, Jean-Marie Jacquet, David R. Gilbert, Sebastian Hunt
CP3
1989 Specifying Concurrent Systems Using Logic
David R. Gilbert
FORTE1
1988 A LOTOS to PARLOG Translator
David R. Gilbert
FORTE1