Anil Wipat

dblp:08/2183 · DBLP profile ↗
← Back
28ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-7310-4191ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 4 since 2021Systems, architecture and hardware · 5
YearPublicationVenuePosition
2024 NeDRex-Web: An Interactive Web Tool for Drug Repurposing by Exploring Heterogeneous Molecular Networks
abstract
Finding new indications for approved drugs is a promising alternative to the often very lengthy and expensive process of de novo drug development. Systems medicine has brought forth several different approaches to tackle this important task. We recently published NeDRex, a network medicine tool for the identification of disease modules and drug repurposing. NeDRex-Web (https://web.nedrex.net) brings existing and new features of the NeDRex platform to a user-friendly and research-oriented web application, enabling online exploration of large heterogeneous molecular networks. Focusing mainly on drug repurposing, NeDRex-Web implements customizable disease module identification and drug prioritization workflows to support users of diverse backgrounds in their research. Users are assisted during every step of their analysis, including the definition of relevant input sets, the selection from various algorithms for module identification or drug prioritization, and the prioritization of the results by their statistical significance. A guided connectivity search provides an easy way to identify links between node sets of interest and can be used to create user-specific induced networks.
Andreas Maier 0009, Mahdie Rafiei, Elisa Anastasi, Olga I. Zolotareva, James Skelton, Maria L. Elkjaer, Ana I. Casas, Cristian Nogales, Harald H. H. W. Schmidt, Tim Kacprowski, David B. Blumenthal, Anil Wipat, Sepideh Sadegh, Jan Baumbach
BIBM12
2022 Automatic Diverse Subset Selection From Enzyme Families by Solving the Maximum Diversity Problem
abstract
Enzymes are being increasingly exploited in various industries for their potential as biocatalysts. Increasing the portfolio of available and useful biocatalysts depends on the reliable annotation of enzyme catalytic function. However, the required quality of such annotation can only be confidently guaranteed through experimental characterisation in the laboratory. The selection of catalytically diverse enzyme panels for experimentally characterisation is therefore an important step for shedding light on the currently unannotated proteins in enzyme families. As current selection methods lack efficiency and scalability, and are non systematic, we present a novel approach for the automatic selection of subsets from enzyme families. A tabu search algorithm solving the maximum diversity problem for sequence identity was designed and implemented, and applied on three diverse enzyme families. We show that this approach automatically selects panels of enzymes that contain high richness and relative abundance of the known catalytic functions, and outperforms other methods such as k-medoids.
Christian Atallah, Katherine James, Zhen Ou, James Skelton, David Markham, James Finnigan, Simon Charnock, Anil Wipat
CIBCB8
2022 Integration of probabilistic functional networks without an external Gold Standard
abstract
BACKGROUND: Probabilistic functional integrated networks (PFINs) are designed to aid our understanding of cellular biology and can be used to generate testable hypotheses about protein function. PFINs are generally created by scoring the quality of interaction datasets against a Gold Standard dataset, usually chosen from a separate high-quality data source, prior to their integration. Use of an external Gold Standard has several drawbacks, including data redundancy, data loss and the need for identifier mapping, which can complicate the network build and impact on PFIN performance. Additionally, there typically are no Gold Standard data for non-model organisms. RESULTS: We describe the development of an integration technique, ssNet, that scores and integrates both high-throughput and low-throughout data from a single source database in a consistent manner without the need for an external Gold Standard dataset. Using data from Saccharomyces cerevisiae we show that ssNet is easier and faster, overcoming the challenges of data redundancy, Gold Standard bias and ID mapping. In addition ssNet results in less loss of data and produces a more complete network. CONCLUSIONS: The ssNet method allows PFINs to be built successfully from a single database, while producing comparable network performance to networks scored using an external Gold Standard source and with reduced data loss.
Katherine James, Aoesha Alsobhe, Simon J. Cockell, Anil Wipat, Matthew R. Pocock
BMC Bioinform.4
2021 Modelling The Fitness Landscapes of a SCRaMbLEd Yeast Genome
abstract
The use of microorganisms for the production of industrially important compounds and enzymes is becoming increasingly important. Eukaryotes have been less widely used than prokaryotes in biotechnology, because of the complexity of their genomic structure and biology. The Yeast2.0 project is an international effort to engineer the yeast Saccharomyces cerevisiae to make it easy to manipulate, and to generate random variants using a system called SCRaMbLE. SCRaMbLE relies on artificial evolution in vitro to identify useful variants, an approach which is time consuming and expensive. We developed an in silico simulator for the SCRaMbLE system, using an evolutionary computing approach, which can be used to investigate and optimize the fitness landscape of the system. We applied the system to the investigation of the fitness landscape of one of the S. cerevisiae chromosomes, and found that our results fitted well with those previously published. Our simulator can be applied to the analysis of the fitness landscapes of any organism for which SCRaMbLE has been implemented.
Bill Yang, Goksel Misirli, Anil Wipat, Jennifer Hallinan
CIBCB3
2019 Harmonizing semantic annotations for computational models in biology
abstract
Life science researchers use computational models to articulate and test hypotheses about the behavior of biological systems. Semantic annotation is a critical component for enhancing the interoperability and reusability of such models as well as for the integration of the data needed for model parameterization and validation. Encoded as machine-readable links to knowledge resource terms, semantic annotations describe the computational or biological meaning of what models and data represent. These annotations help researchers find and repurpose models, accelerate model composition and enable knowledge integration across model repositories and experimental data stores. However, realizing the potential benefits of semantic annotation requires the development of model annotation standards that adhere to a community-based annotation protocol. Without such standards, tool developers must account for a variety of annotation formats and approaches, a situation that can become prohibitively cumbersome and which can defeat the purpose of linking model elements to controlled knowledge resource terms. Currently, no consensus protocol for semantic annotation exists among the larger biological modeling community. Here, we report on the landscape of current annotation practices among the COmputational Modeling in BIology NEtwork community and provide a set of recommendations for building a consensus approach to semantic annotation.
Maxwell Lewis Neal, Matthias König 0003, David P. Nickerson, Goksel Misirli, Reza Kalbasi, Andreas Dräger, Koray Atalag, Vijayalakshmi Chelliah, Mike T. Cooling, Daniel L. Cook, Sharon M. Crook, Miguel de Alba, Samuel H. Friedman, Alan Garny, John H. Gennari, Padraig Gleeson, Martin Golebiewski, Michael Hucka, Nick S. Juty, Chris J. Myers, Brett G. Olivier, Herbert M. Sauro, Martin Scharm, Jacky L. Snoep, Vasundra Touré, Anil Wipat, Olaf Wolkenhauer, Dagmar Waltemath
Briefings Bioinform.26
2016 Computational intelligence for metabolic pathway design: Application to the pentose phosphate pathway
abstract
Metabolic engineering is increasingly being used for the production of industrial products such as pharmaceuticals and enzymes. These chemicals have traditionally been chemically synthesized, but the application of synthetic biology techniques to microbes facilitates faster, cheaper production. Modelling and the integration of existing data can help inform the design of synthetic pathways. We applied an evolutionary algorithm to a flux balance model of metabolism in the industrially important bacterium Bacillus subtilis. Our target metabolites are sedoheptulose-7-phosphate and riboflavin, components of the pentose phosphate pathway. The algorithm combines the results of the flux balance analysis with phylogenetic information derived from data warehouses, to predict several potential interventions to the metabolic network, mostly involving knockouts of genes related to the pathway.
James Skelton, Jennifer Hallinan, Anil Wipat
CIBCB4
2016 Engineering bacterial populations for pattern formation
abstract
The automated design of synthetic biological circuits is an active area of research. A particularly promising area of research is the engineering of populations of communicating bacteria, in order to produce behaviour more complex than is possible with the engineering of individual bacteria. We present a computational approach to the engineering of communicating bacterial populations, using a multi-level approach. Circuits are designed using an evolutionary algorithm, at a high level of abstraction, with an agent-based model. Evolved agents can then be mapped onto previously-defined, lower-level components such as Standard Virtual Parts. This approach is applied to the evolution of a two-dimensional pattern, the French Flag.
Daniel Sutantyo, Christopher Walker, Nicholas deBono, Jarryd Vargas, Anil Wipat, Jennifer Hallinan
CIBCB5
2016 Annotation of rule-based models with formal semantics to enable creation, analysis, reuse and visualization
abstract
MOTIVATION: Biological systems are complex and challenging to model and therefore model reuse is highly desirable. To promote model reuse, models should include both information about the specifics of simulations and the underlying biology in the form of metadata. The availability of computationally tractable metadata is especially important for the effective automated interpretation and processing of models. Metadata are typically represented as machine-readable annotations which enhance programmatic access to information about models. Rule-based languages have emerged as a modelling framework to represent the complexity of biological systems. Annotation approaches have been widely used for reaction-based formalisms such as SBML. However, rule-based languages still lack a rich annotation framework to add semantic information, such as machine-readable descriptions, to the components of a model. RESULTS: We present an annotation framework and guidelines for annotating rule-based models, encoded in the commonly used Kappa and BioNetGen languages. We adapt widely adopted annotation approaches to rule-based models. We initially propose a syntax to store machine-readable annotations and describe a mapping between rule-based modelling entities, such as agents and rules, and their annotations. We then describe an ontology to both annotate these models and capture the information contained therein, and demonstrate annotating these models using examples. Finally, we present a proof of concept tool for extracting annotations from a model that can be queried and analyzed in a uniform way. The uniform representation of the annotations can be used to facilitate the creation, analysis, reuse and visualization of rule-based models. Although examples are given, using specific implementations the proposed techniques can be applied to rule-based models in general. AVAILABILITY AND IMPLEMENTATION: The annotation ontology for rule-based models can be found at http://purl.org/rbm/rbmo The krdf tool and associated executable examples are available at http://purl.org/rbm/rbmo/krdf CONTACT: [email protected] or [email protected].
Goksel Misirli, Matteo Cavaliere, William Waites, Matthew R. Pocock, Curtis Madsen, Owen Gilfellon, Ricardo Honorato-Zimmer, Paolo Zuliani, Vincent Danos, Anil Wipat
Bioinform.10
2014 Tuning receiver characteristics in bacterial quorum communication: An evolutionary approach using standard virtual biological parts
abstract
Populations of bacteria acting in collaboration can produce complex behaviors which are not achievable by individual cells. There has, consequently, been considerable interest in the engineering of bacterial populations. Here we describe an approach for the engineering of aspects of bacterial quorum communication, using Standard Virtual Parts, a synthetic biology programming language, dubbed SVPWrite, and an evolutionary algorithm. We apply this system to engineering the output characteristics of the subtilin receiver system of Bacillus subtilis. Simple modifications, such as altering the strength of the output response to a subtilin input, are easily achieved. More complex adaptations, such as modifying the shape of the receiver response curve, necessitate alterations to the topology of the regulatory network. More generally, the use of Standard Virtual Parts and a programming language allow circuit design, simulation and evaluation to easily be automated, permitting exploration of a far larger proportion of design space than would be possible using standard manual design approaches.
Jennifer Hallinan, Owen Gilfellon, Goksel Misirli, Anil Wipat
CIBCB4
2014 Composable Modular Models for Synthetic Biology
abstract
Modelling and computational simulation are crucial for the large-scale engineering of biological circuits since they allow the system under design to be simulated prior to implementation in vivo . To support automated, model-driven design it is desirable that in silico models are modular, composable and use standard formats. The synthetic biology design process typically involves the composition of genetic circuits from individual parts. At the most basic level, these parts are representations of genetic features such as promoters, ribosome binding sites (RBSs), and coding sequences (CDSs). However, it is also desirable to model the biological molecules and behaviour that arise when these parts are combined in vivo . Modular models of parts can be composed and their associated systems simulated, facilitating the process of model-centred design. The availability of databases of modular models is essential to support software tools used in the model-driven design process. In this article, we present an approach to support the development of composable, modular models for synthetic biology, termed Standard Virtual Parts. We then describe a programmatically accessible and publicly available database of these models to allow their use by computational design tools.
Goksel Misirli, Jennifer Hallinan, Anil Wipat
ACM J. Emerg. Technol. Comput. Syst.3
2014 Introduction to the Special Issue on Computational Synthetic Biology
abstract
The goal of this special issue is to introduce the field of computational synthetic biology to engineers and computer scientists. The first article gives an introduction to the key biological principles and experimental techniques that support synthetic biology, and it draws analogies with the computing field. This issue also includes five original research articles in computational synthetic biology. The first research article discusses how standards can be used to modularize the design process for genetic circuits. The next two articles introduce new abstraction techniques to improve the efficiency of analysis of genetic circuit models. The last two articles introduce new design techniques that help decouple design from construction. We hope this sampling from the field will help to motivate others to join this exciting and rich area of research.
Chris J. Myers, Herbert M. Sauro, Anil Wipat
ACM J. Emerg. Technol. Comput. Syst.3
2012 Data mining the human gut microbiota for therapeutic targets
abstract
It is well known that microbes have an intricate role in human health and disease. However, targeted strategies for modulating human health through the modification of either human-associated microbial communities or associated human-host targets have yet to be realized. New knowledge about the role of microbial communities in the microbiota of the gastrointestinal tract (GIT) and their collective genomes, the GIT microbiome, in chronic diseases opens new opportunities for therapeutic interventions. GIT microbiota participation in drug metabolism is a further pharmaceutical consideration. In this review, we discuss how computational methods could lead to a systems-level understanding of the global physiology of the host-microbiota superorganism in health and disease. Such knowledge will provide a platform for the identification and development of new therapeutic strategies for chronic diseases possibly involving microbial as well as human-host targets that improve upon existing probiotics, prebiotics or antibiotics. In addition, integrative bioinformatics analysis will further our understanding of the microbial biotransformation of exogenous compounds or xenobiotics, which could lead to safer and more efficacious drugs.
Matthew Collison, Robert P. Hirt, Anil Wipat, Sirintra Nakjang, Philippe Sanseau, James R. Brown
Briefings Bioinform.3
2012 Bayesian integration of networks without gold standards
abstract
MOTIVATION: Biological experiments give insight into networks of processes inside a cell, but are subject to error and uncertainty. However, due to the overlap between the large number of experiments reported in public databases it is possible to assess the chances of individual observations being correct. In order to do so, existing methods rely on high-quality 'gold standard' reference networks, but such reference networks are not always available. RESULTS: We present a novel algorithm for computing the probability of network interactions that operates without gold standard reference data. We show that our algorithm outperforms existing gold standard-based methods. Finally, we apply the new algorithm to a large collection of genetic interaction and protein-protein interaction experiments. AVAILABILITY: The integrated dataset and a reference implementation of the algorithm as a plug-in for the Ondex data integration framework are available for download at http://bio-nexus.ncl.ac.uk/projects/nogold/
Jochen Weile, Katherine James, Jennifer Hallinan, Simon J. Cockell, Phillip Lord, Anil Wipat, Darren J. Wilkinson
Bioinform.6
2011 Model annotation for synthetic biology: automating model to nucleotide sequence conversion
abstract
MOTIVATION: The need for the automated computational design of genetic circuits is becoming increasingly apparent with the advent of ever more complex and ambitious synthetic biology projects. Currently, most circuits are designed through the assembly of models of individual parts such as promoters, ribosome binding sites and coding sequences. These low level models are combined to produce a dynamic model of a larger device that exhibits a desired behaviour. The larger model then acts as a blueprint for physical implementation at the DNA level. However, the conversion of models of complex genetic circuits into DNA sequences is a non-trivial undertaking due to the complexity of mapping the model parts to their physical manifestation. Automating this process is further hampered by the lack of computationally tractable information in most models. RESULTS: We describe a method for automatically generating DNA sequences from dynamic models implemented in CellML and Systems Biology Markup Language (SBML). We also identify the metadata needed to annotate models to facilitate automated conversion, and propose and demonstrate a method for the markup of these models using RDF. Our algorithm has been implemented in a software tool called MoSeC. AVAILABILITY: The software is available from the authors' web site http://research.ncl.ac.uk/synthetic_biology/downloads.html.
Goksel Misirli, Jennifer Hallinan, Tommy Yu, James R. Lawson, Sarala M. Wimalaratne, Mike T. Cooling, Anil Wipat
Bioinform.7
2011 Customizable views on semantically integrated networks for systems biology
abstract
MOTIVATION: The rise of high-throughput technologies in the post-genomic era has led to the production of large amounts of biological data. Many of these datasets are freely available on the Internet. Making optimal use of these data is a significant challenge for bioinformaticians. Various strategies for integrating data have been proposed to address this challenge. One of the most promising approaches is the development of semantically rich integrated datasets. Although well suited to computational manipulation, such integrated datasets are typically too large and complex for easy visualization and interactive exploration. RESULTS: We have created an integrated dataset for Saccharomyces cerevisiae using the semantic data integration tool Ondex, and have developed a view-based visualization technique that allows for concise graphical representations of the integrated data. The technique was implemented in a plug-in for Cytoscape, called OndexView. We used OndexView to investigate telomere maintenance in S. cerevisiae. AVAILABILITY: The Ondex yeast dataset and the OndexView plug-in for Cytoscape are accessible at http://bsu.ncl.ac.uk/ondexview.
Jochen Weile, Matthew R. Pocock, Simon J. Cockell, Phillip Lord, James M. Dewar, Eva-Maria Holstein, Darren J. Wilkinson, David A. Lydall, Jennifer Hallinan, Anil Wipat
Bioinform.10
2010 Standard virtual biological parts: a repository of modular modeling components for synthetic biology
abstract
MOTIVATION: Fabrication of synthetic biological systems is greatly enhanced by incorporating engineering design principles and techniques such as computer-aided design. To this end, the ongoing standardization of biological parts presents an opportunity to develop libraries of standard virtual parts in the form of mathematical models that can be combined to inform system design. RESULTS: We present an online Repository, populated with a collection of standardized models that can readily be recombined to model different biological systems using the inherent modularity support of the CellML 1.1 model exchange format. The applicability of this approach is demonstrated by modeling gold-medal winning iGEM machines. AVAILABILITY AND IMPLEMENTATION: The Repository is available online as part of http://models.cellml.org. We hope to stimulate the worldwide community to reuse and extend the models therein, and contribute to the Repository of Standard Virtual Parts thus founded. Systems Model architecture information for the Systems Model described here, along with an additional example and a tutorial, is also available as Supplementary information. The example Systems Model from this manuscript can be found at http://models.cellml.org/workspace/bugbuster. The Template models used in the example can be found at http://models.cellml.org/workspace/SVP_Templates200906.
Mike T. Cooling, V. Rouilly, Goksel Misirli, James R. Lawson, Tommy Yu, Jennifer Hallinan, Anil Wipat
Bioinform.7
2009 Clustering incorporating shortest paths identifies relevant modules in functional interaction networks
abstract
Many biological systems can be modeled as networks. Hence, network analysis is of increasing importance to systems biology. We describe an evolutionary algorithm for selecting clusters of nodes within a large network based upon network topology together with a measure of the relevance of nodes to a set of independently identified genes of interest. We apply the algorithm to a previously published integrated functional network of yeast genes, using a set of query genes derived from a whole genome screen of yeast strains with a mutation in a telomere uncapping gene. We find that the algorithm identifies biologically plausible clusters of genes which are related to the cell cycle, and which contain interactions not previously identified as potentially important. We conclude that the algorithm is valuable for the querying of complex networks, and the generation of biological hypotheses.
Jennifer Hallinan, Matthew R. Pocock, Stephen G. Addinall, David A. Lydall, Anil Wipat
CIBCB5
2009 Saint: a lightweight integration environment for model annotation
abstract
Abstract Summary: Saint is a web application which provides a lightweight annotation integration environment for quantitative biological models. The system enables modellers to rapidly mark up models with biological information derived from a range of data sources. Availability and Implementation: Saint is freely available for use on the web at http://www.cisban.ac.uk/saint. The web application is implemented in Google Web Toolkit and Tomcat, with all major browsers supported. The Java source code is freely available for download at http://saint-annotate.sourceforge.net. The Saint web server requires an installation of libSBML and has been tested on Linux (32-bit Ubuntu 8.10 and 9.04). Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Allyson L. Lister, Matthew R. Pocock, Morgan L. Taschuk, Anil Wipat
Bioinform.4
2008 Network motifs in context: An exploration of the evolution of oscillatory dynamics in transcriptional networks
abstract
The concept of a network motif-a small set of interacting genes which produce a predictable behaviour at the network level-has attracted considerable attention amongst network analysts. It is of particular interest to synthetic biology, a new discipline which aims to apply engineering principles to biological systems. The modular nature of network motifs would make them ideal candidates for the basic components of an engineered organism. In this paper we investigate the relationship between the presence of network motifs and oscillatory dynamics in a yeast transcriptional network and a set of computational networks, evolved to exhibit oscillatory behaviour. Our results do not support the hypothesis that network motifs are critical to network dynamics, possibly because they are tightly connected to many other components of the complex cell-wide transcriptional network.
Jennifer Hallinan, Anil Wipat
CIBCB2
2007 Motifs and Modules in Fractured Functional Yeast Networks
abstract
The integration of diverse data sets into probabilistic functional networks is an active and important area of research in systems biology. In this paper we fracture a previously published integrated network into its component networks, and investigate the overlap between the information provided by each data set to the final network. Using three-node network motifs as a surrogate for information about genetic circuits, we find that the same motifs are over-represented in all of the networks, but different genes contribute to the motifs in different data sets. We conclude that the data integration approach is valuable because it clearly does combine different insights into a biological system. However, the fact that the information contained in different data sets is so diverse raises issues of how best to perform data integration so as to accurately estimate error rates for different data sets, whilst including as much data as possible in the integrated network
Jennifer Hallinan, Anil Wipat
CIBCB2
2007 Qualitatively modelling and analysing genetic regulatory networks: a Petri net approach
abstract
MOTIVATION: New developments in post-genomic technology now provide researchers with the data necessary to study regulatory processes in a holistic fashion at multiple levels of biological organization. One of the major challenges for the biologist is to integrate and interpret these vast data resources to gain a greater understanding of the structure and function of the molecular processes that mediate adaptive and cell cycle driven changes in gene expression. In order to achieve this biologists require new tools and techniques to allow pathway related data to be modelled and analysed as network structures, providing valuable insights which can then be validated and investigated in the laboratory. RESULTS: We propose a new technique for constructing and analysing qualitative models of genetic regulatory networks based on the Petri net formalism. We take as our starting point the Boolean network approach of treating genes as binary switches and develop a new Petri net model which uses logic minimization to automate the construction of compact qualitative models. Our approach addresses the shortcomings of Boolean networks by providing access to the wide range of existing Petri net analysis techniques and by using non-determinism to cope with incomplete and inconsistent data. The ideas we present are illustrated by a case study in which the genetic regulatory network controlling sporulation in the bacterium Bacillus subtilis is modelled and analysed. AVAILABILITY: The Petri net model construction tool and the data files for the B. subtilis sporulation case study are available at http://bioinf.ncl.ac.uk/gnapn.
L. Jason Steggles, Richard Banks, Oliver Shaw, Anil Wipat
Bioinform.4
2007 Exploring Microbial Genome Sequences to Identify Protein Families on the Grid
abstract
The analysis of microbial genome sequences can identify protein families that provide potential drug targets for new antibiotics. With the rapid accumulation of newly sequenced genomes, this analysis has become a computationally intensive and data-intensive problem. This paper describes the development of a Web-service-enabled, component-based, architecture to support the large-scale comparative analysis of complete microbial genome sequences and the subsequent identification of orthologues and protein families (Microbase). The system is coordinated through the use of Web-service-based notifications and integrates distributed computing resources together with genomic databases to realize all-against-all comparisons for a large volume of genome sequences and to present the data in a computationally amenable format through a Web service interface. We demonstrate the use of the system in searching for orthologues and candidate protein families, which ultimately could lead to the identification of potential therapeutic targets.
Anil Wipat, Matthew R. Pocock, P. A. Lee, Keith Flanagan, J. T. Worthington
IEEE Trans. Inf. Technol. Biomed.2
2006 Clustering and Cross-talk in a Yeast Functional Interaction Network
abstract
Many different clustering algorithms have been applied to biological networks, with varying degrees of success. The output of a clustering algorithm may be hard to interpret in biological terms because such networks are often large and highly interconnected, with structural and functional modules overlapping to varying degrees. In this paper we describe an evolutionary network clustering algorithm specifically designed for the analysis of large, complex biological networks. It identifies variably sized, overlapping clusters of nodes. The identification of points of overlap between clusters facilitates the analysis of the biological nature of crosstalk between functional units in the network. We apply two variants of the algorithm (one using probabilistic weights on edges and one ignoring them) to a recently published network of functional gene interactions in the yeast Saccharomyces cerevisiae and assess the biological validity of the resulting clusters in terms of ontological similarity
Jennifer Hallinan, Anil Wipat
CIBCB2
2006 Taverna: lessons in creating a workflow environment for the life sciences
abstract
Abstract Life sciences research is based on individuals, often with diverse skills, assembled into research groups. These groups use their specialist expertise to address scientific problems. The in silico experiments undertaken by these research groups can be represented as workflows involving the co‐ordinated use of analysis programs and information repositories that may be globally distributed. With regards to Grid computing, the requirements relate to the sharing of analysis and information resources rather than sharing computational power. The myGrid project has developed the Taverna Workbench for the composition and execution of workflows for the life sciences community. This experience paper describes lessons learnt during the development of Taverna. A common theme is the importance of understanding how workflows fit into the scientists' experimental context. The lessons reflect an evolving understanding of life scientists' requirements on a workflow environment, which is relevant to other areas of data intensive and exploratory science. Copyright © 2005 John Wiley & Sons, Ltd.
Thomas M. Oinn, Robert Mark Greenwood, Matthew Addis, Mahmut Nedim Alpdemir, Justin Ferris, Kevin Glover, Carole A. Goble, Antoon Goderis, Duncan Hull, Darren Marvin, Peter Li, Phillip Lord, Matthew R. Pocock, Martin Senger, Robert Stevens 0001, Anil Wipat, Chris Wroe
Concurr. Comput. Pract. Exp.16
2005 A grid-based system for microbial genome comparison and analysis
abstract
Genome comparison and analysis can reveal the structures and junctions of genome sequences of different species. As more genomes are sequenced, genomic data sources are rapidly increasing such that their analysis is beyond the processing capabilities of most research institutes. The grid is a powerful solution to support large-scale genomic data processing and genome analysis. This paper presents the Microbase project that is developing a grid-based system for genome comparison and analysis, and discusses the first implementation of the system (called MicrobaseLite). MicrobaseLite uses a scalable computing environment to support computationally intensive microbial genome comparison and analysis, employing state-of-the-art technologies of Web services, notification, comparative genomics and parallel computing. Microbase will support not only system-defined genome comparison and analysis but also user-defined, remotely conceived genome analysis.
Anil Wipat, Matthew R. Pocock, Pete A. Lee, Paul Watson 0001, Keith Flanagan, James T. Worthington
CCGRID2
2004 Taverna: a tool for the composition and enactment of bioinformatics workflows
abstract
MOTIVATION: In silico experiments in bioinformatics involve the co-ordinated use of computational tools and information repositories. A growing number of these resources are being made available with programmatic access in the form of Web services. Bioinformatics scientists will need to orchestrate these Web services in workflows as part of their analyses. RESULTS: The Taverna project has developed a tool for the composition and enactment of bioinformatics workflows for the life sciences community. The tool includes a workbench application which provides a graphical user interface for the composition of workflows. These workflows are written in a new language called the simple conceptual unified flow language (Scufl), where by each step within a workflow represents one atomic task. Two examples are used to illustrate the ease by which in silico experiments can be represented as Scufl workflows using the workbench application.
Thomas M. Oinn, Matthew Addis, Justin Ferris, Darren Marvin, Martin Senger, Robert Mark Greenwood, Tim J. Carver, Kevin Glover, Matthew R. Pocock, Anil Wipat, Peter Li
Bioinform.10
2004 SARGE: a tool for creation of putative genetic networks
abstract
SUMMARY: SARGE is a tool for creating, visualizing and manipulating a putative genetic network from time series microarray data. The tool assigns potential edges through time-lagged correlation, incorporates a clustering mechanism, an interactive visual graph representation and employs simulated annealing for network optimization. AVAILABILITY: The application is available as a .jar file from http://www.bioinformatics.cs.ncl.ac.uk/sarge/index.html.
O. J. Shaw, Colin Harwood, L. Jason Steggles, Anil Wipat
Bioinform.4
2003 On the Use of Agents in BioInformatics Grid
abstract
My Grid is an e-Science Grid project that aims to help biologists and bioinformaticians to perform workflow-based in silico experiments, and help them to automate the management of such workflows through personalisation, notification of change and publication of experiments. In this paper, we describe the architecture of my Grid and how it will be used by the scientist. We then show how my Grid can benefit from agents technologies. We have identified three key uses of agent technologies in my Grid: user agents, able to customize and personalise data, agent communication languages offering a generic and portable communication medium, and negotiation allowing multiple distributed entities to reach service level agreements.
Luc Moreau 0001, Simon Miles, Carole A. Goble, Robert Mark Greenwood, Vijay Dialani, Matthew Addis, Mahmut Nedim Alpdemir, Rich Cawley, David De Roure, Justin Ferris, Robert J. Gaizauskas, Kevin Glover, Christopher Greenhalgh, Peter Li, Phillip Lord, Michael Luck, Darren Marvin, Thomas M. Oinn, Norman W. Paton, Steve Pettifer, Milena Radenkovic 0001, Angus Roberts, Alan J. Robinson, Tom Rodden, Martin Senger, Nick Sharman, Robert Stevens 0001, Brian Warboys, Anil Wipat, Chris Wroe
CCGRID30